Programming8 min read
Claude Code: stop burning tokens on the wrong effort level
On FrontierCode, Sonnet 5.5 scores 52.1% at xhigh and 46.2% at max, and the max run costs 13 times more. What effort changes in Claude Code and how to set it.
Filip SedivySoftware Engineer · AI/ML & Computer Vision


If your Claude Code session runs on /effort max because max sounds like the safe choice, on some tasks you pay more for a worse result. On FrontierCode 1.1, one of the benchmarks on Anthropic's Sonnet 5.5 launch page, Sonnet 5.5 scores 52.1% at xhigh and 46.2% at max. A task at max costs about $20.78, at xhigh about $1.59.
Most of what gets repeated about effort does not hold up against Anthropic's own documentation. ultrathink does not switch you to max, and effort does not set a thinking budget.
What effort controls#
Effort is a parameter of the request. According to the effort documentation, it affects every output token in the response: the text, the arguments of tool calls, and the thinking when thinking runs. Anthropic calls it a behavioral signal. There is no token budget behind it, the only hard limit is still max_tokens.
What changes is the way Claude works. At a lower level it makes fewer tool calls, merges operations into one call, starts acting without explaining the plan and confirms in a sentence. At a higher level it makes more calls, explains the plan before it acts, writes a detailed summary of the changes and leaves more comments in the code. Anthropic's post on choosing a model and effort level in Claude Code names what you notice in a session: how many files Claude reads, how much it verifies, and how far it pushes through a multi-step task before it checks in with you.
The same post shows one prompt at two levels, where the high-effort path produces roughly seven times more tokens. It is an illustration, not a measurement, so read it as an order of magnitude.


The model sets what Claude is able to do. Effort sets how much of it Claude does on one task. Anthropic splits the diagnosis the same way. If Claude skipped a file, didn't run the tests or left a refactor half done, raise effort. If it worked hard and was still wrong, switch the model. Before either, check the context Claude had, starting with CLAUDE.md and the prompt itself.
The level you are on without knowing it#
There are five levels: low, medium, high, xhigh and max. The scale is calibrated per model, so medium on Opus 5.5 is not the same amount of work as medium on Opus 5.
| Level | What the effort docs say | When the Claude Code docs suggest it |
|---|---|---|
low | Most efficient. Significant token savings with some capability reduction. | Quick exchanges where you review each result: brainstorming, a first sketch, a rename. |
medium | Balanced, with moderate token savings. | Day-to-day work with a clear scope, such as implementing a new feature. |
high | Spends as many tokens as the task needs. | Work where verification matters or edge cases are likely, such as a bug in an existing codebase. |
xhigh | Long-horizon work over 30 minutes, with token budgets in the millions. | Deeper reasoning at a higher token spend. |
max | No constraint on token spending. | Hard problems Claude should work through without you, such as finding security vulnerabilities. Prone to overthinking, so test it first. |
The default depends on the model, and for Sonnet 5.5 also on where you run it. In Claude Code and the Claude apps, Sonnet 5.5 starts at medium, through the API at high. Opus 5.5 starts at medium everywhere, Opus 4.7 at xhigh in Claude Code, and every other model that supports effort at high.
Pick a model and where you run it to see its default effort level. Opus 5.5 starts at medium in Claude Code and in the API. Sonnet 5.5 starts at medium in Claude Code and the Claude apps and at high in the API. Opus 4.7 starts at xhigh in Claude Code and at high in the API. Every other model that supports effort starts at high. Opus 4.6 and Sonnet 4.6 have no xhigh.
This is where an upgrade costs you without you noticing. Opus 5 defaulted to high. According to the Claude Code model configuration docs, Opus 5.5 at medium matches or exceeds Opus 5 at high on Anthropic's coding and knowledge-work evaluations, and at the same level Opus 5.5 thinks more per turn than Opus 5. If you moved to Opus 5.5 and set high because that is what you ran before, you pay for a level the new model does not need. The docs say it directly: start at medium and don't carry the old level over.
Not every model has every level. Opus 4.6 and Sonnet 4.6 have no xhigh, and Claude Code runs a level the model does not support as the next one down, so xhigh on Opus 4.6 runs as high.
Max is not the top of the chart#
Anthropic publishes the benchmarks for the 5.5 models at every effort level, with the cost per task next to the score. The costs below are read off the charts on the Sonnet 5.5 launch page, so take the cents as approximate.
Score against cost per task at each effort level for Sonnet 5.5 and Opus 5.5 on three benchmarks. FrontierCode 1.1: Sonnet 5.5 rises to 52.1% at xhigh for $1.59 and falls to 46.2% at max for $20.78, Opus 5.5 peaks at 54.6% at medium for $0.80. CursorBench 4.0: Opus 5.5 scores 56.0% at high for $3.97 and 57.8% at max for $13.43. Terminal-Bench 4.0: Sonnet 5.5 reaches 70.6% at max for $12.54 per attempt, Opus 5.5 peaks at 66.4% at xhigh.
On FrontierCode 1.1, Sonnet 5.5 climbs from 29.3% at low to 52.1% at xhigh, then falls to 46.2% at max, while the cost per task goes from $1.59 to $20.78. Anthropic explains the drop in a footnote: at max, Sonnet more often ran a code-review skill with many subagents, and those runs ended in a timeout or with edits outside the scope of the task. Opus 5.5 peaks on the same benchmark at medium, 54.6% for $0.80, which is its default.
Max is not always worse. On Terminal-Bench 4.0 it gives Sonnet 5.5 its best score, 70.6% against 61.5% at xhigh, and on CursorBench 4.0 it adds almost eight points over high. It is never cheap, and the price grows faster than the score.


Across these six runs the step from high to max costs between 2.9 and 49.5 times more per task, and it brings anything from minus 3.2 to plus 27.6 points. Until you measure it on your own kind of work, you don't know which end of that range you are on. That is why the docs call max prone to overthinking and tell you to test it before you adopt it broadly.
What it does to time and to mistakes#
The time you wait grows with effort too. In Using Claude Code: Spending your effort, Anthropic's Claude Code team ran the same tasks on Opus 5.5 at each level and wrote down how long they took. A loosely specified fitness app took 1.5 minutes at low, 4 at medium, 11 at high and 67 at max. A task with a detailed spec took 16, 22, 33 and 79 minutes.


The more useful measurement in the same post is about failures. Fable 5.1 ran 370 Terminal-Bench 3.0 attempts at low and at max. At max it solved 214 instead of 140, and the attempts where it missed a case dropped from 59 to 24. The attempts where it picked the wrong reading of the task went up, from 25 to 47. The median attempt used 222k tokens instead of 73k.


The Claude Code docs describe the same thing in other words: at a higher level Claude tested more edge cases, verified more of its work and made more choices on its own. My reading of the numbers is that higher effort makes Claude more thorough about the interpretation it already chose. If that interpretation is wrong, max carries out the wrong thing more carefully, for three times the tokens. A better prompt fixes that, or a short interview before the implementation. Effort does not.
Setting it in Claude Code#
Claude Code takes the first level that applies, from the top:
The environment variable is the only source a skill's frontmatter cannot override, and maxEffortLevel in managed settings is a ceiling nothing can raise. effortLevel in managed settings is only a starting point, users can still change it with /effort.
/effort # opens the slider: Enter saves for this model, s for this session only
/effort high # sets and saves the level for the current model
/effort auto # clears the saved level, back to the model defaultclaude --effort xhigh # this session only
export CLAUDE_CODE_EFFORT_LEVEL=max # the only way to keep max beyond one sessionClaude Code saves the level per model under modelSettings in ~/.claude/settings.json, so each model keeps its own. Two details are easy to miss. Max is never saved: it lasts for the session unless it comes from CLAUDE_CODE_EFFORT_LEVEL, and neither modelSettings nor effortLevel accepts it. And the old top-level effortLevel in your user settings, the form /effort wrote before levels were saved per model, does not count for Opus 5.5 at all. If you set "effortLevel": "high" months ago and the session header on Opus 5.5 says medium, that is why.
ultrathink and ultracode#
Write ultrathink anywhere in a prompt and Claude Code adds an in-context instruction to reason more on that one turn. The effort level sent to the API stays the same. Phrases like "think hard" or "think more" reach the model as ordinary text, Claude Code does not treat them as keywords. That makes ultrathink the cheap way to handle the one hard turn in a session: stay on medium for the routine work and put ultrathink into the prompt about a race condition.
Ultracode is a Claude Code setting, not an effort level. With it on, Claude plans a dynamic workflow with subagents for each substantive task, at whatever level the session runs. Since v2.1.284, /effort ultracode and the Tab toggle in the slider leave the effort level alone. Only claude --effort ultracode also sets it to xhigh. You turn it off with /effort ultracode off, and it is not available when workflows are turned off or the model has no xhigh. Left on for routine work, it means a workflow and the tokens of all its agents for every task Claude considers substantive.
The cache, if you call the API#
In the API, changing the top-level effort in the middle of a conversation invalidates the prompt cache, and the whole history is read again at the full input price. The effort docs describe a beta that avoids it. With the header anthropic-beta: mid-conversation-output-config-2026-07-01 you add a system message with empty content and output_config.effort, and everything before it stays cached. It works on Fable 5.1, Mythos 5.1, Opus 5.5, Opus 5 and Sonnet 5.5. In Claude Code, /effort shows a cache warning when the change would invalidate the cache.
So which level#
Anthropic's advice in the model and effort post is to keep the model's default for most tasks and treat effort as a general preference, not as a decision before every prompt. On Opus 5.5, and on Sonnet 5.5 in Claude Code, that means medium, and leaving it there until a session gives you a concrete reason to change it.
The effort post on claude.dev splits one task across levels instead of picking one for all of it. Write a spec, let Claude interview you about it, implement at low, review the result yourself, then run the verification at high.