What does Extended Thinking actually do, and how is it different from a normal reply?
Extended thinking is the mechanism where Claude performs a round of internal reasoning before producing its final reply. The official documentation frames this as self-assessment — the model first judges whether this particular request genuinely needs extra reasoning to answer well; if it decides it doesn't, that turn may contain no thinking Block at all. If it decides it does, the thinking process unfolds first, and the final answer is produced based on it.
The biggest difference from a normal reply is whether it happens at all every time. Extended thinking isn't a feature that fires on every single turn — it's decided dynamically based on request complexity. Within the same conversation, one turn might include a thinking block while the next, answering a simple question, has none at all. That variation in depth is expected behavior by design, not instability or a malfunction.
What's the actual difference between the legacy budget_tokens and the current effort parameter?
budget_tokens is the older configuration style used on older models, and it's manual: you directly specify a hard Token cap on how much the thinking process can use. The official documentation specifically notes that changing the budget_tokens value between requests invalidates prompt caching breakpoints.
The effort parameter is what current models use instead, set under output_config.effort, offering five levels: low, medium, high, xhigh, and max. The fundamental difference from budget_tokens is that effort is soft guidance — it tells the model roughly which direction to lean its resource allocation, rather than a hard-coded ceiling. Even set to max, there's no guarantee thinking fires on every request; even set to low, the model may still decide thinking is needed for a genuinely complex request. One thing that carries over unchanged from budget_tokens: changing the effort level between requests still invalidates cache breakpoints — that cost didn't go away just because the parameter changed.
How do I decide which effort level fits my application?
The official documentation lays out a specific comparison: max means the most thinking at the greatest depth, with no constraint on thinking length; xhigh thinks more readily and at greater depth than high; high is the default for most models, firing on most requests that would benefit from it; medium is the default for Claude Opus 5.5, offering moderate thinking that may skip simple queries entirely; low minimizes thinking and prioritizes response speed.
In practice, the choice comes down to the nature of the task: simple, latency-sensitive, high-volume repeated calls are a good fit for low, or simply not setting it explicitly and relying on the model's default. Tasks needing multi-step reasoning, code debugging, or complex analysis usually pay off better at high or above. The thing actually worth watching is the other half of cost control — max_tokens is the hard cap on total output (thinking plus reply combined), while effort only determines roughly what share of that cap goes toward thinking. Both parameters need to be set together for a complete picture.
How is Extended Thinking actually billed, and does the thinking content shown on screen match what's actually charged?
Not exactly. Billing is based on the full thinking Token count recorded in the usage.output_tokens_details.thinking_tokens field, and that number gets billed as ordinary output tokens. But the thinking content actually displayed to you may be a summarized version, not a verbatim rendering of the model's complete internal reasoning. In other words, the length of thinking text you see on screen can't be used to directly estimate how much a given request actually cost — what actually matters is the returned thinking_tokens number.
This gap matters for cost control: if you judge spend purely by eyeballing how long the displayed thinking content looks, it's easy to underestimate actual cost, since the full internal reasoning can run considerably longer than the summary shown on screen.
The official documentation's billing example makes this concrete: a request's returned usage shows input_tokens: 25 and output_tokens: 348, with output_tokens_details.thinking_tokens recording 312 — meaning that of those 348 output tokens, 312 were actually consumed by the thinking process, leaving only 36 tokens as the actual reply content shown to the user. For billing purposes, all 348 tokens count as output tokens together, with no separate pricing between thinking and the reply.
Extended thinking can meaningfully improve answer quality on tasks that genuinely need multi-step reasoning, and effort's soft guidance is more flexible than the old hard cap, removing the need to hand-tune token counts for every task type; the cost is that the thinking process itself counts toward output token billing, and changing the effort level invalidates prompt caching breakpoints — a cost that needs separate evaluation for high-frequency, latency-sensitive applications.