Bible Network Crypto DeFi Onchain RWA AI Agent Stablecoin Chain SAFU CryptoTax DeFAI AGI Claude Me Claude Skill Claude Design Claude Cowork
Independent Media
Not affiliated with any project
Learn Claude Skills. Do Everything Better.
claudeskill-me.com
LATEST
Claude Suddenly Stopped Using Your Skills After /compact? It's Not Broken — It's Designed Not to Restore Them Automatically  ·  Claude Code Now Supports AGENTS.md — But Half the Articles Online Describe the Old Rules, and CLAUDE.local.md Can Silently Break It  ·  Your Hook Says "Blocking Error" But the File Still Changed? PostToolUse and PreToolUse Don't Actually Block the Same Way  ·  Enabled Auto Mode and Thought You Were Safe? Permission Modes and Sandbox Boundaries Are Two Different Layers — Conflating Them Is How Things Break  ·  Messages API Adds On-Demand Compaction in Beta: Developers Decide When to Compact, Not the System  ·  Claude Cowork and Chat Officially Merge, Launching Claude Docs and Claude Slides Alongside It
Glossary · Prompt Engineering

Extended Thinking

Prompt Engineering intermediate

30-Second Version · For the impatient
The internal reasoning Claude performs before replying — the model decides per request whether and how deeply to think, now steered by the effort parameter rather than the legacy hard-capped budget_tokens.
Full Explanation +
01 · What is this?

What does Extended Thinking actually do, and how is it different from a normal reply?

Extended thinking is the mechanism where Claude performs a round of internal reasoning before producing its final reply. The official documentation frames this as self-assessment — the model first judges whether this particular request genuinely needs extra reasoning to answer well; if it decides it doesn't, that turn may contain no thinking Block at all. If it decides it does, the thinking process unfolds first, and the final answer is produced based on it.

The biggest difference from a normal reply is whether it happens at all every time. Extended thinking isn't a feature that fires on every single turn — it's decided dynamically based on request complexity. Within the same conversation, one turn might include a thinking block while the next, answering a simple question, has none at all. That variation in depth is expected behavior by design, not instability or a malfunction.

02 · Why does it exist?

What's the actual difference between the legacy budget_tokens and the current effort parameter?

budget_tokens is the older configuration style used on older models, and it's manual: you directly specify a hard Token cap on how much the thinking process can use. The official documentation specifically notes that changing the budget_tokens value between requests invalidates prompt caching breakpoints.

The effort parameter is what current models use instead, set under output_config.effort, offering five levels: low, medium, high, xhigh, and max. The fundamental difference from budget_tokens is that effort is soft guidance — it tells the model roughly which direction to lean its resource allocation, rather than a hard-coded ceiling. Even set to max, there's no guarantee thinking fires on every request; even set to low, the model may still decide thinking is needed for a genuinely complex request. One thing that carries over unchanged from budget_tokens: changing the effort level between requests still invalidates cache breakpoints — that cost didn't go away just because the parameter changed.

03 · How does it affect your decisions?

How do I decide which effort level fits my application?

The official documentation lays out a specific comparison: max means the most thinking at the greatest depth, with no constraint on thinking length; xhigh thinks more readily and at greater depth than high; high is the default for most models, firing on most requests that would benefit from it; medium is the default for Claude Opus 5.5, offering moderate thinking that may skip simple queries entirely; low minimizes thinking and prioritizes response speed.

In practice, the choice comes down to the nature of the task: simple, latency-sensitive, high-volume repeated calls are a good fit for low, or simply not setting it explicitly and relying on the model's default. Tasks needing multi-step reasoning, code debugging, or complex analysis usually pay off better at high or above. The thing actually worth watching is the other half of cost control — max_tokens is the hard cap on total output (thinking plus reply combined), while effort only determines roughly what share of that cap goes toward thinking. Both parameters need to be set together for a complete picture.

04 · What should you do?

How is Extended Thinking actually billed, and does the thinking content shown on screen match what's actually charged?

Not exactly. Billing is based on the full thinking Token count recorded in the usage.output_tokens_details.thinking_tokens field, and that number gets billed as ordinary output tokens. But the thinking content actually displayed to you may be a summarized version, not a verbatim rendering of the model's complete internal reasoning. In other words, the length of thinking text you see on screen can't be used to directly estimate how much a given request actually cost — what actually matters is the returned thinking_tokens number.

This gap matters for cost control: if you judge spend purely by eyeballing how long the displayed thinking content looks, it's easy to underestimate actual cost, since the full internal reasoning can run considerably longer than the summary shown on screen.

Sources: Effort - Claude Platform Docs, Steering thinking and cost - Claude Platform Docs
Real-World Example +

The official documentation's billing example makes this concrete: a request's returned usage shows input_tokens: 25 and output_tokens: 348, with output_tokens_details.thinking_tokens recording 312 — meaning that of those 348 output tokens, 312 were actually consumed by the thinking process, leaving only 36 tokens as the actual reply content shown to the user. For billing purposes, all 348 tokens count as output tokens together, with no separate pricing between thinking and the reply.

Common Misconceptions +
✕ Misconception 1
× Misconception: setting effort to max guarantees a full thinking process fires on every request, when actually: effort is soft guidance on allocation direction, not a hard guarantee — the model still decides for itself whether a given request genuinely needs thinking; max only means "if it decides to think, it can think as deeply and as long as possible," not "it must think every time"
✕ Misconception 2
× Misconception: the length of thinking content shown on screen can be used to estimate a request's actual cost, when actually: the displayed thinking content may be a summarized version, while actual billing is based on the returned thinking_tokens number — the two don't necessarily match, and judging by displayed length alone tends to underestimate real cost
The Missing Link +
Direct Impact

Extended thinking can meaningfully improve answer quality on tasks that genuinely need multi-step reasoning, and effort's soft guidance is more flexible than the old hard cap, removing the need to hand-tune token counts for every task type; the cost is that the thinking process itself counts toward output token billing, and changing the effort level invalidates prompt caching breakpoints — a cost that needs separate evaluation for high-frequency, latency-sensitive applications.

Ask a Question
Please enter at least 10 characters
More Related Topics