All research

The System That Says No

14 pages · 12 cited sources

Open the paperDownload

What this paper finds

  • In a controlled simulation with perfect data, a material-balance plot fitted at R² = 0.9998 while the resulting estimate was 8.2% too high — and overestimated original oil in place by +160%, +90% and +50% at 3%, 7% and 20% depletion.
  • On a benchmark built to measure knowing-when-not-to-answer, the highest-accuracy model scored 0.81 on accuracy but only 0.46 on abstention. The two abilities are close to decoupled.
  • Reasoning fine-tuning degrades abstention by 24% on average — even in the maths and science domains those models are explicitly trained on.