OpenAI's CFO Proposes 'Useful Intelligence Per Dollar' to Replace Cost-Per-Token as the AI ROI Metric

OpenAI CFO Sarah Friar published a framework on July 17 arguing that cost-per-token is the wrong way to judge enterprise AI spend, proposing “Useful Intelligence per Dollar” instead. The framework rests on four questions: how much useful work a system actually completes — issues resolved, code shipped, contracts reviewed; the true cost per successful task, counting employee review time and rework rather than token price alone; dependability, tracked as results that are ready to use, need correction, or need escalation; and whether returns improve as compute and infrastructure compound at scale. Friar cites GPT-5.6 Sol as the efficiency proof point: a new state of the art on the Artificial Analysis Coding Agent Index using 54% fewer output tokens than a competing model, and 72.7% accuracy on the DeepSWE v1.1 benchmark against a rival’s 69.9% at 36.2% lower estimated API cost.

The metric Friar wants replaced is the one OpenAI itself set three weeks earlier. GPT-5.6’s June preview priced the family at $5 input / $30 output per million tokens for the flagship Sol, $2.50/$15 for the mid-tier Terra, and $1/$6 for the fastest Luna tier — the exact per-token numbers “Useful Intelligence per Dollar” argues buyers should stop anchoring on. That preview also detailed why the model matters beyond price: Sol set a new state of the art on Terminal-Bench 2.1 and showed a large enough jump in offensive-security capability that OpenAI paired the release with over 700,000 A100-equivalent GPU hours of red-teaming and, at the US government’s request, initially restricted access to a small group of trusted partners rather than shipping broadly.

Friar’s four questions are a genuinely useful checklist for any team building an ROI case internally or for a client — task completion, true cost, dependability, and scale economics beat a bare token price every time. But it’s worth naming what it is: a vendor CFO defining the yardstick her own company’s product should be measured by, published the same month that product’s list price shipped in three tiers. The framework is worth adopting. The specific numbers behind it are worth re-running against a client’s own task data before they go in a deck.