GPT-6 Astra Reviewed 41 Legal Documents in Minutes — and Caught Every Planted Error
GPT-6 Astra completed a 41-document financial-statement tie-out for legal-AI platform Legora in a single continuous run, catching all four errors OpenAI’s team had deliberately planted in the test set — including a £500,000 gap concealed inside a revenue note. OpenAI reports that’s roughly a 40% performance improvement over the prior model generation on the same workflow. Legora runs this kind of document review for more than 100,000 professionals across over 1,800 in-house legal departments and law firms in 50-plus markets, which makes the benchmark less a lab demo than a proxy for what’s about to change in document-heavy professional work: due diligence, audit support, contract review, the parts of consulting and legal engagements that consume the most billable hours for the least differentiated judgment.
The catch-rate is the part worth sitting with. A single missed figure in a financial-statement tie-out is the kind of error that survives a rushed human review and shows up later as a client-relationship problem. A model that finds a six-figure discrepancy buried in a revenue note, every time, changes what “reviewed” means as a deliverable — not faster review of the same quality, but a different quality bar at the same speed.
OpenAI’s “Path to Astra” disclosure is why this shipped on the timeline it did. Two days before Astra’s release, OpenAI disclosed that the model was the first of its models to cross the “Critical” cybersecurity capability threshold under its own Preparedness Framework — plausible offensive capability comparable to a skilled human red-teamer finding zero-days. The model’s safety overview confirms OpenAI deliberately delayed parts of the rollout by several weeks specifically to harden protections against cyber misuse before wider release, and states Astra is the most-aligned model the company has shipped to date.
For firms evaluating frontier models as vendors rather than just tools, that sequencing is itself useful information: a lab naming a specific capability threshold before general release, rather than burying it in a post-release system card, is a rare enough disclosure pattern to weigh alongside the benchmark numbers. The Legora result is a real, citable proof point that document-review throughput gains are showing up now, in production, at the volume law firms and consulting due-diligence teams actually operate at — not as a roadmap promise.