Claude Fable 5.1 Costs Up to 45% Less on Agentic Workloads, and Anthropic Says It Beats GPT-5.6 Sol

Claude Fable 5.1 costs up to 45% less to run on heavily agentic workloads than its predecessor, driven by a 75% cut in cache-read pricing to $0.25 per million tokens — even as Anthropic claims the model outperforms Fable 5, Opus 5, and GPT-5.6 Sol across a slate of new benchmarks. Fable 5.1, generally available, scores 52.6% on Terminal-Bench-Science versus Fable 5’s 24.7%, and 73.4% on CursorBench 3.2.0; a restricted-access sibling, Claude Mythos 5.1, hits 60.9% on Terminal-Bench 4.0 and is gated behind new Cyber and Life Sciences Verification Programs limited initially to US organizations. Standard API rates land at $10 per million input tokens and $50 per million output tokens. Anthropic is tightening safety tooling alongside the capability gains: cybersecurity-use safeguards now produce 60% fewer false positives, biology-safeguard activations on benign requests dropped 85%, and Fable 5.1 can be used to discover software vulnerabilities but explicitly not to build exploits for them.

The price move lands inside a pattern, not in isolation. OpenAI cut GPT-5.6 API prices by as much as 80% the same week Fable 5.1 shipped — Luna, its cheapest tier, now runs $0.20 per million input tokens — and credited part of the drop to GPT-5.6 Sol rewriting its own production inference kernels to cut serving costs 20%. OpenAI claims Luna beats Claude on the Agents’ Last Exam benchmark at nearly 99% lower cost per task; Anthropic’s own benchmark table claims the opposite result against GPT-5.6 Sol. Both labs are pointing at the same handful of benchmarks and claiming to win them, which is itself the more useful signal than either scorecard: frontier pricing is now falling in lockstep with capability claims, not trading off against them.

For anyone advising a client on model selection, the practical read isn’t “which model wins” — the two labs’ own benchmark claims contradict each other — it’s that the cost of running agentic workloads at frontier quality keeps dropping fast enough that a six-month-old procurement decision is worth revisiting on price alone, independent of whether the underlying capability gap has actually changed.