OpenAI's Astra for Law Beat Web-Search-Only Astra by 15 Points on Legal Research Accuracy

54% versus 38.7%. That’s the correctness gap OpenAI reports between Astra for Law — its GPT-6 Astra model tuned specifically for legal work — and the same base model using only web search, on Vals AI’s Legal Research Bench. Astra for Law pairs the model with a dedicated legal-search index spanning more than 230 million URLs of U.S. case law, statutes, regulations, and administrative decisions, plus a partnership with the Free Law Project’s CourtListener collection covering over 99.9% of published U.S. case law. At the model’s highest reasoning setting, it also surfaced 24% more relevant reference cases than the web-search-only version on case-law questions. OpenAI shipped 26 partner plugins — iManage, Intapp, Relativity, and Clio among them — so firms can run it alongside existing systems instead of migrating work into ChatGPT wholesale.

The product logic tracks two case studies OpenAI has published in the same stretch. Law firm Cooley used GPT-6 Astra to build “GO Public,” a purpose-built agent system that accelerates Form S-1 drafting for IPO clients, and says the underlying architecture extends to other capital-markets transaction types. And on a financial-statement tie-out benchmark, an agent built on GPT-6 Astra reviewed 41 documents in one run and caught all four intentionally planted errors — including a £500,000 gap hidden in a revenue note — a roughly 40% improvement over the prior model generation on that task.

The pattern across all three: measurable accuracy gains on the exact document-heavy, detail-dense, high-stakes work that fills a consulting firm’s due-diligence and regulatory-filing engagements. A single missed figure in a tie-out or a missed precedent in a brief has real financial and legal consequences — which is why vertical tuning, not just a bigger general model, is where the benchmark gains are actually showing up.