Claude Now Has a 'Think Harder' Button. Here's When to Press It — and When Not To.
Anthropic shipped an effort toggle with Claude Opus 5. Google released three tiers of Gemini Flash and deprecated temperature settings. Every major AI provider made the same call this week: match the task to the reasoning depth. Here's how.
By Forge Team · Examples are illustrative composites
The choice that used to happen automatically is now yours to make. Claude Opus 5, released July 24, ships with a three-position effort toggle — low, medium, or high — that controls how much reasoning the model applies before responding. The same week, Google released Gemini 3.6 Flash in three separate tiers and deprecated temperature, top-p, and top-k settings across all future models. Both companies made the same decision in the same week: the controls that matter are now in your hands, and the right setting is something you decide based on what the task actually needs.
What shipped
Opus 5 launched at $5 per million input tokens, $25 per million output — matching the Opus 4.8 price. The difference is in what low, medium, and high effort actually mean. On low, it answers fast. On high, it reasons through the problem, self-verifies its conclusions, and tends to surface issues you didn't ask about. Analyst Simon Willison noted on July 24 that Opus 5 outperforms Anthropic's flagship Fable 5 on several coding and knowledge-work benchmarks at roughly half the cost, and behaves proactively — when it can't access data you need directly, it builds a workaround rather than reporting a dead end.
Gemini 3.6 Flash, released July 21 at $1.50/$7.50 per million tokens with a 1 million token context window, now ships in three tiers: Flash for general use, Flash-Lite for high-volume cheap tasks, and Flash Cyber for security-sensitive work. Google's simultaneous deprecation of temperature, top-p, and top-k removes the technical knobs that gave experienced users fine-grained control over output behavior. The expectation going forward is that you phrase the instruction differently — those adjustments belong in the prompt now.
What changes on Monday
You now have a standing decision to make at the start of every AI task: how much reasoning does this actually need? The costs of getting it wrong run in both directions.
Running everything at high effort adds cost and latency to work that doesn't benefit from deeper reasoning. Running analysis tasks at low effort produces output that sounds thorough but hasn't been genuinely examined — you often won't notice until someone downstream catches the problem.
Three practical guidelines:
Formatting, drafts, routine rewrites → low effort or Flash-Lite. Rephrasing a subject line, generating five tagline options, reformatting a table — the model doesn't need to think hard. Fast and cheap is appropriate.
Analysis, strategy, pressure-testing → high effort or frontier tier. Diagnosing a process problem, evaluating options with real tradeoffs, checking whether a plan has a fatal flaw — these are the tasks where deeper reasoning changes the output in ways that matter.
When unsure, start medium and escalate only if the output is shallow. The cheapest mistake is running a task twice at different settings. The expensive mistake is shipping shallow analysis because it looked complete.
Elena — the retention strategy problem
Elena is head of customer success at a 120-person B2B SaaS company. She uses AI daily: meeting summaries, renewal emails, feedback pattern analysis. For almost all of it, medium or default effort is fine.
Last quarter she used the same settings to write the retention strategy for a cohort that was 34% behind on renewals. The output looked solid — clear sections, data references, a framework. It missed a shared characteristic across the at-risk accounts that would have changed the entire approach. She found it herself, three hours later, while preparing for the VP presentation.
The gap: a well-framed prompt on high effort would have caught the oversight for a few cents more in compute. The real cost was the three hours and the near-miss in a $200k decision.
High effort matters most when the stakes are highest. The rest of the time, medium works.
James — the counterpoint
James is a senior editor at a 12-person strategy consulting firm. After reading about the effort toggle, he switched all his first-draft work to high effort.
The drafts got longer and more hedged. The model spent its reasoning budget on qualifications and caveats rather than clean structure. High-effort reasoning is optimised for problem-solving — it second-guesses, surfaces edge cases, checks for gaps. That's exactly wrong for a first draft, where you want confident forward momentum.
He switched back to medium for drafts and now reserves high effort for briefings where he needs the model to evaluate an argument before he builds from it.
The one thing
The effort toggle isn't a performance upgrade — it's a calibration tool. Anthropic and Google both moved in the same direction this week — one with a named toggle, one by removing the technical alternatives. The professionals who use these controls deliberately, rather than leaving everything on auto, will get better results at lower cost. That decision happens before you hit send.
<BlogPracticeSection />Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Forge turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Your AI Just Got an Effort Dial. Here's When to Turn It Up.
Claude Opus 4.8 shipped with explicit effort controls — users choose whether the model reasons carefully or answers fast, at 3x the cost difference. Here's how to decide which tasks deserve which mode.
4 min readYour AI Vendor Just Drew an Ethical Line. Here's Why That Affects Your Workflow.
Court documents unsealed July 2 showed Anthropic drew two non-negotiable redlines — no mass surveillance, no autonomous weapons — and lost government access because of it. The 19-day Fable 5 suspension was the operational consequence. Here's what it means for how you build with AI.
5 min readYour AI Provider Now Needs Government Permission to Ship. Here's What That Means for Your Workflows.
Both OpenAI and Anthropic have their best models gated by government approval — not pricing tiers, not waitlists. If you built workflows around those models, here are three things to do before it happens again.
5 min read