OpenAI Cut GPT-5.6 API Prices Up to 80% — Partly by Having the Model Optimize Its Own Infrastructure

OpenAI cut GPT-5.6 API pricing starting July 30: Luna, the fastest and cheapest model in the family, dropped 80% to $0.20 per million input tokens and $1.20 per million output tokens; Terra, the balanced everyday-work model, dropped 20% to $2 and $12 per million tokens; Sol’s pricing held steady. A new Fast mode replaced Priority Processing, running Sol up to 2.5x faster than standard processing at twice the price with no change in intelligence.

The more interesting detail is where the savings came from. OpenAI credits part of the cut to GPT-5.6 Sol optimizing its own production stack — autonomously rewriting inference kernels to cut end-to-end serving costs 20%, and improving speculative-decoding efficiency by more than 15%. A companion engineering post on the same release adds the specifics: Sol rewrote OpenAI’s production GPU kernels directly in Triton/Gluon, ran hundreds of self-directed experiments to improve its own draft model, and — per OpenAI’s own benchmark claim — now beats Anthropic’s Claude Fable 5 on the Artificial Analysis Coding Agent Index at under half the cost.

OpenAI’s own framing: Luna now matches year-old frontier-class model performance at roughly 6 cents on the dollar and nearly nine times the speed, and undercuts Claude on the Agents’ Last Exam benchmark at an estimated cost per task nearly 99% lower.

For consulting engagements building a cost model around agentic workloads, the useful number here isn’t “AI got cheaper” — it’s the specific per-token deltas: an 80% cut on the tier most high-volume agent workloads would actually run on. That’s a real input for a build-vs-buy calculation, not a marketing line.