Microsoft's New Coding Model Costs a Quarter as Much as Its Predecessor and Still Scores Higher

Microsoft’s MAI-Code-1.1-Flash, now in production in GitHub Copilot and VS Code, produces higher-quality code at a quarter of the cost of its predecessor, MAI-Code-1.0, which launched in June. Microsoft states the new model streams 25% faster while using 25% fewer tokens to complete a given task — a straight cost-per-token improvement, not just a raw speed gain, on tooling developers already use daily.

The benchmark numbers back the cost claim with quality data rather than asking buyers to take the efficiency gain on faith: a 22% improvement on Terminal-Bench 2.1 inside GitHub Copilot CLI, a 15% improvement on .NET-focused tasks, a 4% increase in code survival rates, and a 9% increase in developer return visits. Microsoft says the model was refined through “hundreds of thousands of reinforcement-learning environments in GitHub Copilot,” and that the team prioritized command-line and .NET performance specifically because that’s what developer feedback pointed to — summarized internally as “ship, learn, improve, repeat.”

For companies evaluating coding-assistant spend, this is a concrete data point in a pattern that’s been showing up across AI vendors all year: iteration cycles compressing cost per output without asking teams to sacrifice quality to get there. A model that costs a quarter as much and scores better on the benchmarks that map to real developer workflows changes the math on tool selection and renewal conversations — not because “AI is getting cheaper” in the abstract, but because a specific successor model, shipped roughly two months after its predecessor, quantifies exactly how much cheaper and where the gains show up.

The model is available now for testing in GitHub Copilot, with Microsoft soliciting feedback through GitHub issues rather than treating the release as finished. For consulting engagements scoping developer-tooling decisions, the useful takeaway isn’t the specific 25% figure — it’s the cadence: coding-model economics are moving fast enough that a tool evaluation done six months ago is already stale.