Grok 4.6 Scored Highest on a Biosecurity Refusal Benchmark — By Balancing Refusal Against Usefulness

xAI published independent third-party evaluation results for Grok 4.6 on two biosecurity benchmarks, run by LatchBio. On BioSecBench-Refusal — which tests whether a model can distinguish legitimate dual-use biological research requests from disguised hazardous ones across BSL-1 through BSL-4 scenarios — Grok 4.6 scored a trial-weighted harmonic mean of 62.1%, the highest of all models tested, and was the only model to score above 50% on both underlying measures simultaneously: refusing 59.2% of red-team hazardous-intent tasks while still completing 64.8% of routine, dual-use-adjacent legitimate research tasks.

That dual performance is the finding worth naming. Models routinely optimize for one measure at the other’s expense — either over-refusing legitimate research or under-refusing genuinely hazardous requests — and most published benchmarks report refusal rates in isolation rather than the tradeoff. On the companion BioSecBench-Surveillance benchmark, which measures usefulness for detecting and characterizing emerging pathogens, Grok 4.6 averaged 53.5%, behind the model xAI refers to as “Opus 5” but ahead of “GPT-5.6 Sol.”

xAI describes a layered safeguard approach — refusal training, inference-time safeguards, behavioral controls, post-deployment monitoring — built explicitly to avoid degrading legitimate research use. Anthropic has been running the same tradeoff from the other direction: an August update to Claude Fable 5’s biology safeguards rewrote the constitutional guidelines behind its safety classifiers specifically to cut false positives, producing roughly an 85% reduction in biology-related fallbacks to a less-capable model across product surfaces — a 67% drop on Claude.ai alone — while keeping restrictions on professional virology, toxicology, and drug-development queries intact.

Two labs, two different benchmark methodologies, the same underlying admission: refusal alone isn’t a safety metric worth publishing on its own anymore. The number that matters is how much legitimate work a safeguard costs — and both companies are now putting that number in writing.