GPT-5.6 Beats Claude Fable 5 by 13 Points on Agents' Last Exam, at Half the Coding-Agent Cost

OpenAI moved GPT-5.6 to general availability on July 9, 2026, and its flagship tier, Sol, scored 53.6 on Agents’ Last Exam — 13.1 points ahead of Claude Fable 5 — while setting a new state of the art of 80 on the Artificial Analysis Coding Agent Index using less than half the tokens, time, and cost of competing models. Sol also posted 92.2% on BrowseComp and 62.6% on OSWorld 2.0. Pricing lands at $5 per million input tokens and $30 per million output tokens for Sol, with two cheaper tiers — Terra at $2.50/$15 and Luna at $1/$6 — sharing the same 1-million-token context window.

The launch lands in a month where every frontier lab has been racing on price as much as capability. xAI’s Grok 4.5, which shipped the day before, undercuts GPT-5.6 outright at $2 per million input tokens and $6 per million output, and xAI claims it resolves tasks using an average of 15,954 output tokens against Anthropic’s Opus 4.8 (max) at 67,020 — a 4.2x efficiency gap Elon Musk framed as “an Opus-class model, but faster, more token-efficient and lower cost.” Claude Sonnet 5, which launched ten days earlier at $2/$10 introductory pricing, made the same argument from Anthropic’s side: agentic capability approaching Opus-class performance at a fraction of the price.

For companies architecting an AI stack rather than picking a single vendor, the pattern across three launches in ten days is the story: benchmark leadership and pricing are both moving fast enough that a model choice made in June is already a live cost-optimization question in July. Sol’s coding-agent efficiency claim — half the tokens, time, and cost for a new state-of-the-art score — is the kind of number worth verifying against a real workload before it becomes the default, not taking on a vendor’s word alone.