Skip to content
Tutorials

How to Use Effort Control in Opus 5 and Cut Your Bill Without Losing Quality

The new dial in Claude Opus 5 lets you set how hard the model works on each request. Used well it changes your unit economics more than any model swap. A practical guide to when low is enough and when maximum earns its price.

2 min read
Abstract illustration of two people working with a neon green knot and purple machinery, symbolizing collaboration in complex systems.

The most consequential change in Claude Opus 5 is not on any leaderboard. It is effort control: a setting from low to maximum that decides how much work the model puts into an answer. Instead of paying for deliberation everywhere, you pay for it where it changes the outcome.

Why this beats picking a cheaper model

Until now the choice was coarse: run an expensive smart model on everything, or a cheap weaker one on everything. The dial breaks that into per-task decisions. The published numbers show the shape of it: Opus 5 scores 61 on the Intelligence Index at max effort and 59 at high. Those two points can cost a multiple of the price, and in most everyday work they change nothing you would notice.

Where low effort is enough

  • Summarising and rewriting text you already have.
  • Classification, tagging, structured extraction from invoices, emails, tables.
  • Translation and copy editing.
  • Bulk operations across many records, where unit cost dominates.

Anthropic reports that even on its lowest effort setting, Opus 5 passes more tasks on the Zapier automation benchmark than competing models do. That makes low your sensible default for anything that runs in a loop.

Where maximum earns its price

  • Debugging and refactoring code you did not write.
  • Finding contradictions across long documents: contracts, policies, specifications.
  • Multi-step work where an error in step three quietly poisons everything after it.
  • Analysis where one answer is correct and plausible is not good enough.

Speed is a separate axis

Alongside effort there is the fast serving mode: about 2.5 times default speed for double the price. It earns its keep when a human or an interface is waiting for the response, and it is money burned on anything that runs in the background.

A rule of thumb to start with

Begin at a middle setting and walk down until quality starts to bother you. Most teams discover they were paying maximum effort for work that a setting two steps lower handles identically. At $5 per million input tokens and $25 per million output, that difference shows up in your invoice within a week, not a year.

author avatar
Promptyze
Promptyze covers generative AI in plain English — hands-on reviews, tutorials and daily news, fact-checked and hype-free.

Promptyze

ADMINISTRATOR

Promptyze covers generative AI in plain English — hands-on reviews, tutorials and daily news, fact-checked and hype-free.

$ sitemap --all The whole site in one place — so you never get lost.