Mistral's Fortran Migration Found the Failure Mode on Both Ends of Agent Autonomy
Give agents full autonomy on a hard legacy-code migration and you get code that technically compiles but is “Fortran retyped in C++” — not actually modernized. That’s the finding from a Mistral AI case study published September 9, describing a project to migrate 40,000 of 300,000 total lines of a physics-intensive reservoir simulator, originally written in Fortran 77 for a European energy operator, over to C++. The codebase had no existing test suite and scattered documentation, so before touching any code Mistral’s team built a numerical parity harness, used a custom parser to map caller-callee trees, and ran Mistral OCR against legacy PDF documentation to reconstruct what the system was actually supposed to do. Only then did they deploy Mistral’s Vibe CLI to spawn multi-agent teams — planner, coder, tester, and reviewer roles — migrating the code module by module in self-contained subtrees under roughly 10,000 lines each, with human review gates between phases.
The team’s finding cuts against both extremes: full agent autonomy produced syntactically-translated-but-not-modernized code, while pure manual migration stalled on the codebase’s harder bugs. The hybrid approach — structured human supervision at defined checkpoints — was what actually worked, and Mistral distilled it into three principles: establish parity verification before migration starts, complete documentation before deploying any agents, and build the review-gate structure in from the beginning rather than bolting it on.
It’s a useful counterpoint to how Anthropic described its own internal migrations two months earlier — Bun co-founder Jarred Sumner used Claude Code to port roughly a million lines from Zig to Rust in under two weeks, with 100% of Bun’s existing test suite passing before merge. The difference is instructive: Sumner’s migration had a comprehensive existing test suite to check agent output against, while Mistral’s Fortran codebase had none — which is exactly why full autonomy failed for Mistral but scaled cleanly for Anthropic. The lesson generalizes past either vendor: agent autonomy is only as safe as the verification harness underneath it, and legacy scientific code without tests needs the human gates that a well-tested codebase can mostly skip.