Anthropic's Middle Model Now Cites a Customer's Own Benchmark as Proof
Anthropic priced Claude Opus 5 at $5 per million input tokens and $25 per million output — identical to Opus 4.8, the model it replaces in that tier — while claiming performance close to Fable 5, the more expensive flagship above it. On Anthropic’s own “CursorBench 3.2” figure, Opus 5 lands within 0.5% of Fable 5’s peak score at half the cost; on ARC-AGI 3 it scores three times higher than the next-best model; on OSWorld 2.0 it beats Fable 5’s best result at a third of the price. It also ships adjustable “effort” settings for trading intelligence against token spend, plus self-verification loops that iterate on a task before returning it. Anthropic reports the lowest misaligned-behavior score (2.3) among its recent releases, while keeping Opus 5 intentionally behind Fable-class systems on cybersecurity-exploitation and advanced-biology benchmarks as a guardrail.
The CursorBench name isn’t incidental. Cursor built that benchmark specifically because engineer Nate Schmidt stopped trusting public leaderboards to predict how a model would handle messy, underspecified real prompts — a bare stack trace with the word “fix,” a task seeded with a deliberately wrong hint. Fable 5 scored 72.9% there in July, a result Schmidt called a qualitative shift: he stopped re-supplying context and started handing off ambiguous, gnarly problems outright. Opus 5’s benchmark table is, in effect, citing a customer’s own distrust of vendor marketing as its proof point.
The pricing move extends a pattern. Sonnet 5 launched a month earlier at introductory pricing under Opus 4.8, explicitly framed as “approaching Opus 4.8 performance at a fraction of the price.” Opus 5 repeats that move one tier up — same price as the model it replaces, closing the gap to the model above it.
For consulting engagements advising on model selection, the practical read is narrower than “which model is smartest.” The cost case for defaulting to the priciest tier keeps eroding every few months, and CursorBench-style internal evals — not vendor benchmark tables — are what should settle the question for a given workload.