Grok 4.5 Says It's Opus-Class for a Fraction of the Price
Grok 4.5 launched July 8 with a specific claim built into the marketing: xAI CEO Elon Musk called it “an Opus-class model, but faster, more token-efficient and lower cost.” The pricing backs that up on paper — $2 per million input tokens and $6 per million output tokens — and xAI cites an efficiency figure of 15,954 average output tokens to resolve a task, 4.2 times fewer than Anthropic’s Opus 4.8 (max) at 67,020. It’s the first Grok model trained alongside Cursor, the coding platform xAI agreed to acquire, and xAI is pointing it at contract review, regulatory analysis, and financial-statement analysis — knowledge work outside its original coding lane.
The pitch — a frontier-class model at a fraction of the cost — is exactly the comparison companies are now equipped to make for themselves. Anthropic’s own guidance on choosing a model and effort level in Claude Code frames the decision as two separate levers: model selection sets the capability ceiling, effort level sets how thoroughly the model checks its own work, and the two shouldn’t be conflated. Applied to Grok’s claim, the real diligence question isn’t “is it cheaper” — it’s whether a 500,000-token-context model tuned for coding and agentic tasks is being asked to do the same job as Opus, at the same effort level, or whether the price comparison is doing work the benchmark isn’t.
That’s the discipline companies evaluating AI vendors keep skipping: matching model tier to task complexity before comparing sticker price. A $2-per-million-token model that needs three retries to hit the same output quality isn’t actually four times cheaper. Grok 4.5’s efficiency numbers are real and worth taking seriously — xAI says its own engineers at Tesla and SpaceX already rely on it in production — but the number worth tracking isn’t cost per token. It’s cost per finished, verified task, which is the metric none of the vendor pricing pages report.