Anthropic’s prompting guide for Opus 5 says: “Claude Opus 5 has a 1M token context window as both the default and the maximum, and its instruction following, tool calling, and reasoning stay consistent throughout the window.” No number backs the second half. So I checked it.
Context Arena, an independent site, runs a test called MRCR: it hides 8 identical requests (say, 8 requests for a poem about the same topic) in a long conversation, then asks for one of the answers, such as the 4th poem. Here are the scores of Opus 5 and of the older Opus 4.8:
| Context length | 8K | 16K | 32K | 64K | 128K | 256K | 512K |
|---|---|---|---|---|---|---|---|
| Opus 5 | 98.7% | 99.9% | 99.3% | 99.9% | 91.3% | 65.6% | 42.5% |
| Opus 4.8 | 97.4% | 96.1% | 93.0% | 90.7% | 75.2% | 61.8% | 39.8% |
There’s no result at 1M. So Opus 5 is much better at 128K, and only a little better after that.
Anthropic used to publish such tests itself: Opus 4.7 scored 32% on MRCR at 1M tokens. Then Anthropic dropped MRCR from Opus 4.8 on. With Opus 5 it also dropped GraphWalks, the test an Anthropic employee had called the better one for long context. Not good! Not good at all! Meanwhile, Anthropic’s own page on context windows says: “As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
Then my own tests on Opus 5 and 5.5: how often did Claude break five of my rules below 150K tokens, and above 300K, in the same sessions?
Three rules showed no clear change, and one had too few cases to tell. The fifth, the cd rule, broke in 17% of the cases below 150K and in 39% above 300K. Refused tool calls stayed under 1% at every size. Above 400K my data gets thin because I clear/compact the context at 310K.
So Anthropic’s promise is clearly inflated. Up to 400K most rules held, one clearly did not, and nothing shows what happens at 1M.
By default, Claude Code compacts a 1M window at about 967K tokens. I start Claude with:
claude --autocompact 343k
Compaction fires about 33K tokens before that number, so my sessions compact at about 310K, close to where the cd rule jumped from 21% to 39%. I don’t go lower, because compaction has a price too: a 2026 study on other models found that compaction kept only 17% to 23% of the rules typed in the chat.
This was one chapter of a whole book
You just read one chapter of Delphi in all its glory – AI-assisted development: more than 400 pages about Claude Code and AI for Delphi programmers. Written by a Delphi programmer, with the failures documented next to the wins.