Switching Models and Effort on the Fly: Claude Code’s /model and /effort

This article is one chapter of my book Delphi in all its glory – AI-assisted development, the fifth book of the series.

Claude Code lets us choose between different models (Opus, Sonnet, Haiku) and effort levels (low, medium, high, xhigh, max). The core idea is simple: not every task deserves the same horsepower. Renaming a variable does not need Opus at max effort. Hunting a subtle memory corruption bug in a 4000-line unit does not deserve Haiku at low effort. Matching the tool to the job saves you time, tokens, and money.

Note: Because Claude is bad at Delphi, I keep the settings way higher than what the internet recommends.

Quick Switching: The Commands

Model switching takes one command:

/model opus
/model sonnet
/model haiku

No need to restart the session, no loss of context. Everything you discussed so far stays in the conversation.

There is also a visual picker. Type /model by itself (no argument) and a menu appears. Use the arrow keys to select a model. While you are in the menu, look at the bottom: there is an effort slider. Use the left and right arrow keys to adjust it. This is a neat shortcut because you can change both model and effort in one go. For effort alone:

/effort low
/effort medium
/effort high
/effort xhigh
/effort max

Same deal. Instant, no restart needed.

There is also Alt+P, which toggles between models without clearing your prompt. Handy when you are halfway through typing a request and realize you want Opus instead of Sonnet.

When to Switch

Here is my practical breakdown for Delphi work:

Opus + Max effort – For the really hard stuff: multi-unit refactoring with tricky dependencies, tracking down a bug that spans several layers of your application, architecture questions, reviewing complex legacy code. Uses significantly more tokens.

Opus + xHigh effort – Good for routine tasks where you already know what you want and just need Claude to type it out: adding a bunch of similar properties, writing repetitive code, reformatting, simple edits across multiple files.

Haiku + Low effort – Quick questions. “What does this compiler switch do?” “What is the difference between TList and TObjectList?” Haiku answers in seconds and costs almost nothing. Switch back to Opus when you need real work done.

The /effort command works mid-conversation, so a typical session might look like this: start on Opus high for the main task, drop to /effort medium for a batch of simple follow-up edits, bump back to /effort high when you hit an unexpected problem. It takes two seconds each time.

Update (bad news)

Under the new Claude Code, you get a warning that the whole conversation will be re-read when you switch the model, or the effort level:

Change effort level?  Your next response will be slower and use more tokens

This conversation is cached for the current effort level. Switching to xhigh means the full history gets re-read on your next message.
▎ 1. Yes, switch to xhigh
▎ 2. No, go back

For a model switch this is real: if the context window is full, every token gets re-processed (and, on a metered plan, re-charged). None of the ways out is good:

  • Clear the current context, then switch the model.

  • Open a new window and work there under the other model – works if this parallel session doesn’t interfere with the current one and you don’t need the current conversation history.

  • Switch anyway, and pay to re-read the whole context.

Whatever you do, it’s bad.

Anthropic’s docs say that on most models each effort level has its own cache4, so the warning is real. Three models are the exception: on Opus 5.5, Sonnet 5.5 and Fable 5.1 (with a Claude subscription or an API key) the cache survives, and Claude Code does not even ask.

A bug report5 says the cache survives an effort switch. It tested xhigh on Sonnet 4.6, a model without xhigh. Claude Code ran that call at high, so the effort never changed: its full cache hit proves nothing.

I had Claude run the experiment on Opus 5, Fable 5.1 and Sonnet 4.6. The result backs the warning, not the bug report: on Opus 5 and Sonnet 4.6, the first message after an effort change processed the whole conversation again. On Fable 5.1 the cache survived.

Switching back to the old level costs less: Opus 5 found its old cache for that level and reused it. The docs keep that cache alive for one hour after its last use on a subscription, five minutes otherwise.

Persistent Settings

If you always want a specific setup, put it in your settings file (settings.json):

{
  "model": "opus",
  "effortLevel": "high"
}

Or use environment variables:

set ANTHROPIC_MODEL=opus
set CLAUDE_CODE_EFFORT_LEVEL=high

These set your defaults. You can still override them with /model and /effort during a session.

What About Auto Effort?

If you have used Cursor, you know it has an “auto” mode where the tool decides the right effort level for each task. Simple question? Low effort. Complex refactoring? High effort. No manual switching needed.

Claude Code does not have this. Not really.

There is /effort auto, but all it does is reset your effort level to the model’s default (medium for Opus 4.6 and Sonnet 4.6). It does not look at your task and decide “this one needs max effort.” It just means “go back to the default.”

Within a given effort level, Claude does have something called Adaptive Thinking Mode. At medium effort, Claude might skip its reasoning step for a trivial question but think deeply about a complex one. So, there is some automatic adjustment happening, but it works within the bounds you set. If you set low effort, Claude will not suddenly switch to deep reasoning just because your question is hard. It stays within its lane.

The practical consequence: you are the one making the judgment call. You know your codebase, you know whether the next task is trivial or complex, and you switch accordingly. For Delphi developers this actually makes sense, because Claude consistently underestimates the difficulty of Delphi tasks. A refactoring that would be straightforward in Python might involve tricky unit dependencies, form file updates, and compiler quirks in Delphi. You are better positioned than any automatic system to know when to crank up the effort.

Until Anthropic adds a true auto-effort feature (and I hope they do), the best approach is to develop the habit of switching. It takes two seconds, and once it becomes muscle memory, you will not even think about it.

The OpusPlan Hybrid

There is one more option worth mentioning: OpusPlan mode. Activate it with:

/model opusplan

This is a hybrid approach. Claude uses Opus (the most capable model) for planning and reasoning, then automatically switches to Sonnet (faster and cheaper) for execution. The idea is that you want the best brain for figuring out what to do, but you do not need the best brain for actually typing out the code once the plan is clear.

In theory, this gives you adaptive model selection without manual switching. In practice, my experience has been mixed. The results were sometimes worse than just using Opus for everything, because the handoff between planning and execution can lose nuance. Your mileage may vary. Try it on a few tasks and see if it works for your workflow.

Plan Expensive, Build Cheap

Today I spent a session with Fable 5, Anthropic’s “big-brain” model, designing a bug-fixing skill for my Delphi products.

Then came the moment to implement the skill: rename a folder, update fifteen cross-references, write two skill files from an approved spec. Mechanical work. And I caught myself: why would I pay big-model prices for find-and-replace?

So, I asked Fable: “Can Sonnet implement your plan? How do I hand it over to your cheaper brother?” The answer was better than I expected.

Two different jobs

A complex task hides two jobs with very different skill requirements.

Design: understand the problem, weigh options, catch the trap you didn’t see. This is where the expensive model earns its money.

Execution: apply the decisions. Rename, edit, wire, verify. A cheaper model does this fine, IF the decisions are already made.

The trick is the handover: make the expensive model write its plan into a file. A complete, self-contained briefing that assumes ZERO shared context.

A briefing that a cheap model can execute cold contains:

  • The goal, in two sentences.

  • Every decision already made. No open questions. None.

  • Exact file paths, verified beforehand (my Fable session grepped every cross-reference and listed them file by file).

  • Text that must be copied verbatim, marked as such.

  • Verification steps with expected results (“rerun this grep, expect zero hits”).

  • The definition of done.

The litmus test: search your plan for “figure out”, “decide later”, “choose the best”. Every one of those is a design decision that will now be made by the cheap brain. Send the plan back to the expensive model until the list is empty.

Three ways to hand it over

1. Fresh session

Just open a second console:

claude --model claude-sonnet-5
> Read MyPlan.md and implement it.

2. Switch models mid-session. Type /model sonnet (or press Alt+P) in the same session and say “go”. You keep the whole conversation, nothing to write down.

But the cheap model now drags your entire design discussion along as context, and long contexts make every model dumber (the flashlight problem, again). Acceptable when the task is small and you’re in the flow. Lazy, works, but not clean!

3. Stay expensive, delegate cheap. The big model stays in charge and spawns sub-agents that run on the cheaper model (an agent file can pin model: sonnet in its header, see the chapter about sub-agents later in this book, see the chapter “Sub-agents”). Most hands-off: the big model briefs the workers, checks their output, and reports back.

You still pay big-model rates for every supervision turn, so it’s the priciest of the three cheap routes. Use it when you’ll be away from the keyboard and want quality control built in.

When NOT to do this

Hand nothing to the cheap model while the task still contains design work. It will not stop and ask. It will improvise, confidently, and you’ll pay twice: once for the wrong implementation, once for the expensive model to untangle it. A plan session plus one correct implementation is cheaper than three wrong ones.

When Sonnet finishes, a short session with the big model reviewing the changes costs little and catches the places where the briefing wasn’t as complete as you thought.


  1. https://code.claude.com/docs/en/prompt-caching#changing-effort-level

    ↩︎

  2. Claude-code bug report #63962, filed 2026-05-30

    ↩︎


This was one chapter of a whole book

You just read one chapter of Delphi in all its glory – AI-assisted development: more than 400 pages about Claude Code and AI for Delphi programmers. Written by a Delphi programmer, with the failures documented next to the wins.

Leave a Comment

Scroll to Top