What this paper finds
- On multimodal tasks, perception accuracy stayed above 99% across modalities while computational accuracy on the same inputs was “often nearing zero” past a computational-load threshold. The model reads correctly, then computes badly.
- On an engineering benchmark spanning nine domains and twenty-seven models, 79.5% of frontier-model errors were arithmetic rather than conceptual — the understanding is largely present, the execution is not.
- In a peer-reviewed 10,000-trial study, task-specific deterministic tools moved accuracy from 11% to 84% and 36% to 95% — outperforming both retrieval augmentation and a general code interpreter.