OpenAI's New GPT-6 Sol Claims Claude-Level Coding Performance at a Fifth of the Cost

OpenAI expanded its GPT-6 family on September 22 with two new models, Sol and Luna, alongside the existing Astra, each priced roughly 50% below their GPT-5.6 predecessors. Sol now runs $2 per million input tokens and $10 per million output tokens, down from $4/$20; Luna drops to $0.10/$0.50 per million tokens from $0.20/$1.20. OpenAI also says Sol makes roughly half as many factual mistakes as the model it replaces.

The benchmark claims matter more than the headline pricing for anyone budgeting agentic-AI spend: OpenAI reports Sol scored 68.8% on the DeepSWE v1.1 coding benchmark at “xhigh” reasoning effort — within 1.1 points of Claude Fable 5’s 69.9% — while costing roughly 80% less per task, and matched Claude Opus 5 on OSWorld 2.0 computer-use tasks at a similar cost reduction. On AutomationBench, a professional-work test, Sol scored 33.2% at $0.27 per task, ahead of Claude Opus 5. These are OpenAI’s own benchmark numbers, not independently verified, but they land in the same week as Anthropic’s own Claude Opus 5.5 launch at 40% lower cost than its predecessor — the second major model-cost repricing from a frontier lab in seven days.

The same-day companion release matters more for anyone running production agents than the headline pricing: an upgraded prompt-caching system now offers up to 90% discounts on cached input tokens reused within a 30-minute window, plus new diagnostics showing exactly where a prompt breaks cache boundaries, and the ability to change reasoning effort mid-task without invalidating the cache. OpenAI frames it explicitly around “persistent agents that work for hours on complex tasks” — refactoring codebases, producing research documents — where token costs previously scaled with context length regardless of how much of that context was actually new.

For companies running side-by-side pilots of Claude and GPT-6 variants, the practical effect isn’t which lab wins a specific benchmark — it’s that frontier-model unit costs are falling fast enough that a pricing model built on per-token estimates from two months ago is probably already stale.