GPT-6 Astra Crosses OpenAI's 'Critical' Cyber Threshold — and Ships Anyway, Slowly
GPT-6 Astra, which OpenAI released on September 3, is the first of the company’s models to cross the “Critical” cybersecurity capability threshold under its own Preparedness Framework — the line at which a model can independently discover and exploit previously unknown vulnerabilities in well-protected systems. Astra hits 100% on ExploitBench and a 42.4% success rate on ExploitGym, up from 30.3% for predecessor GPT-5.6 Sol, and was trained on more than 100,000 GPUs at OpenAI’s Stargate site in Texas — the company’s largest training run yet.
What makes the release notable isn’t just the capability jump. OpenAI flagged the threshold crossing two days early in a standalone post, Path to Astra, and its safety overview for the finished model states OpenAI deliberately delayed parts of the rollout by several weeks specifically to strengthen protections against cyber misuse before wider release. That’s a real-time capability disclosure tied to a disclosed release delay, not a risk assessment buried in a system card after the model is already shipping — a disclosure pattern most frontier labs haven’t matched.
The capability isn’t abstract. In a same-day case study OpenAI published alongside the safety overview, legal-tech platform Legora used Astra to complete a 41-document financial-statement tie-out in a single run, catching all four planted errors — including a £500,000 gap hidden in a revenue note — for roughly a 40% performance improvement over the prior model on that workflow.
For any organization evaluating frontier-model vendors, Astra is a template worth naming out loud: a lab disclosing a specific capability-tier crossing before general release, and pointing to the exact weeks it spent hardening the model in response. Procurement and risk teams sizing up AI vendors should be asking whether their other candidates disclose this way at all, or only after the fact.