All research

The Model Is Not the Product

13 pages · 14 cited sources

Open the paperDownload

What this paper finds

  • On an engineering benchmark spanning nine domains, the best of twenty-seven models reached 65.4% final-answer accuracy but only 42.7% reasoning-trace fidelity — and 79.5% of frontier-model errors were arithmetic rather than conceptual.
  • A published six-agent system improved reservoir history matching by 95% on a toy case, 69% on a moderate one, and 13% on a real field. Performance collapses precisely as problems become real.
  • In a 10,000-trial peer-reviewed study, equipping a model with task-specific deterministic tools moved accuracy from 11% to 84% and from 36% to 95% — outperforming both retrieval augmentation and a general code interpreter.