AI · Opinion
AI labs shipped System 2 and left the hard part of the 2021 IBM paper to you
A 2021 IBM paper on fast and slow AI spent its pages on the arbiter that decides when to think hard. Five years on, the labs sell the slow thinking and hand the deciding to a dropdown.
In October 2021, a year before ChatGPT, ten researchers from IBM, Tulane, Udine and a few other places put a six-page preprint on arXiv called “Thinking fast and slow in AI: the role of metacognition”. It proposed an architecture named SOFAI. A fast solver answers every problem first, from past experience, and reports how confident it is. A slow solver reasons and searches, costs more, and never runs unless something calls it. Most of the paper is about that something: a metacognitive module that decides whether the fast answer is good enough.
The paper is circulating again this week, and the quick verdict on it is that it is obsolete. GPT-5 already knows when to reason, the argument goes, based on a thinking parameter you pass in, like xhigh. I think that reply gets it backwards. The part of SOFAI that the industry shipped is the slow solver, and reasoning models are a very good one. The part it skipped is the module the paper was actually about. The dial you set to xhigh is the metacognition, and you are the one doing it.
Look at what the paper asks its arbiter to know before it spends anything. It keeps a “model of self”: which solver handled this kind of task before, what that cost in time and memory, and what reward the decision earned. It checks whether there are enough resources left to even run the careful check. Then it activates the slow solver only if the extra expected reward beats the expected cost of running it, and the authors say they combine confidence and expected value by multiplying them, a risk-averse choice. You can argue with the details, and in one paragraph they admit they assumed enough past experience on exactly the same problem, which is a big assumption. But every input has a unit and a source.
Now look at what shipped. OpenAI launched GPT-5 in August 2025 with a router that picked between a fast model and a thinking model per request. On launch day it broke. Sam Altman wrote that “the autoswitcher broke and was out of commission for a chunk of the day, and the result was GPT-5 seemed way dumber.” Within days the model picker was back in ChatGPT. The API, meanwhile, went the other way: reasoning effort became a parameter the developer sets, and the levels kept growing up to xhigh. Anthropic added an effort parameter to Claude in late 2025 too. Both designs admit the same thing. Nobody trusted a learned arbiter enough to let it decide alone, so the decision went to a menu.
The cost of that shows up in the other direction as well. A December 2024 paper from Tencent AI Lab researchers, titled “Do NOT Think That Much for 2+3=?”, measured o1-style models on that exact question and found they used 1,953% more tokens than conventional models to reach the same answer. That is a system with a slow solver and no working MC1, the paper’s cheap first-pass check that should have waved the fast answer through.
The obvious objection is that SOFAI was tested on toy problems, and that is fair. The 2025 follow-up in npj Artificial Intelligence ran it on navigation in constrained grids. A second 2025 preprint from the IBM group, SOFAI-LM, paired a plain language model with a reasoning model under a metacognitive monitor on graph coloring and code debugging, and reports matching or beating the standalone reasoning models in accuracy with much lower inference time. I would not bet a product on graph coloring. I would bet that the idea scales better than a dropdown does.
The other objection is that Kahneman’s book has not aged well. That is partly true. Its chapter on priming leaned on studies that failed to replicate, and Kahneman said in 2017 that he had placed too much faith in underpowered studies. None of that touches SOFAI, because the architecture never needed the psychology to be right. It needs one cheap solver, one expensive solver, and a controller that keeps score of which one paid off. Database query planners have done a version of this for decades.
What I would like the labs to ship is the model of self. Log, per task type, how often the fast answer was accepted, how often the slow path changed the answer, and what that cost, then let the arbiter use those numbers and show them to the user. If a lab published a single chart of “slow path changed the answer” rates by effort level, we would learn more about reasoning models than from another benchmark table. Until then, every developer picking xhigh by gut feel is running the 2021 paper’s MC2 step by hand, without the data the authors said it needs.
Prompted by Thinking fast and slow in AI: the role of metacognition, arxiv.org.