
I’ll just say it: What the hell is going on with AI “reasoning”?
Sorry for the air quotes. That punctuational side-eye was more common in 2024, when the specially trained cousins of LLMs now known as “large reasoning models,” or LRMs, were still new. Nowadays it may seem downright churlish, though, given that a “general-purpose reasoning model” from OpenAI solved a famous open mathematical research problem in one shot in May 2026. Still, I’m not sure how else to acknowledge my intellectual whiplash over the scientific interpretation of what these AI systems are actually doing.
Reasoning comes in many technically defined forms, but the basic procedure is easily recognizable: arriving at a sound conclusion by linking together intermediate steps that logically follow from each other. We do this with thoughts; LRMs use so-called chains of thought, a term of art for the streams of synthetic text that the models emit before arriving at an answer to a complex query. One minute, the idea that AI could reason via these chains was being prominently and credibly critiqued (by a team of researchers from Apple) as an “Illusion of Thinking” subject to “complete accuracy collapse” under surprisingly simple conditions. The next minute, LRMs were bagging gold medals at the International Mathematical Olympiad, a feat so challenging that “even very successful mathematicians and scientists may well highlight [it] on their CVs all their lives,” as the scientist and AI critic Gary Marcus and Ernest Davis wrote in 2025. If that’s not a sign of “real” reasoning, what is?
But wait — soon after, more research, from the Santa Fe Institute, showed that LRMs can crush even carefully designed benchmarks for reasoning (like a collection of analogy-like visual puzzles) using mere “surface-level ‘shortcuts.’” What they were doing looked less like generalizable reasoning than just gaming the system. Then, as if on cue, another “hold my beer” moment: Google DeepMind and the mathematician Terence Tao (the GOAT!) used AI to rediscover or improve the solutions to 67 problems “spanning mathematical analysis, combinatorics, geometry, and number theory.” Deal with it, haters!
What about additional evidence that LRMs can’t reason reliably, even when they possess the necessary algorithm and computational budget to do so, and suffer from a list of scientifically documented failure states long enough to use as a Slip ’N Slide? Whatever — I guess that’s just “jagged intelligence” for you (AI-speak for “when it works, it works”).
And so it went from late 2025 into 2026. I’ve been a science journalist for 20 years and an AI journalist for half of that, so I know better than to expect tidy consistency out of rapidly advancing research. But even for me, this back-and-forth has been a bit much. To quote Al Pacino in The Insider, “I’m getting two things: pissed off, and curious.” I don’t believe there’s fraud to be found here. I just want to know which way is up. Can AI reasoning somehow be both BS and not at the same time? And if so, how on Earth does that work?
I knew just who to call first.
![]()
Melanie Mitchell’s career in AI stretches back to the 1980s, but lately she’s earned a reputation as an au courant AI truth teller, penning lucid explainers for Science and her widely read newsletter, as well as conducting research at the Santa Fe Institute. (The study about “surface-level ‘shortcuts’” is hers.) When I asked her what we actually know about AI reasoning, her answer was brief enough to fit on an index card.
“Number one: It works. It improves things,” she said, referring to LRMs’ superior accuracy on reasoning tasks compared to LLMs. “Number two: The actual text that’s generated” — i.e., the chain of thought that every LRM is trained to produce to improve its performance — “isn’t necessarily faithful to what’s going on [inside the model]. And number three: A lot of that text isn’t even useful. You can actually take it out.”
Let’s unpack numbers two and three, because that’s where the superposition of “BS and not” actually lives. Chains of thought were half-discovered, half-devised in 2022 as a prompting hack for LLMs: Provide them with examples of written-out reasoning (or, famously, just ask them to “think step by step”), and they’ll suddenly give less boneheaded answers to simple logic and math problems. LRMs, starting with OpenAI’s o1 model in 2024, are trained to automate this trick by generating such prompts — also called reasoning traces or thinking tokens — and then feeding them back to themselves. Because LRMs are essentially just language models, those extra bits of text create what looks convincingly like a paper trail of the model’s “thought process.”
Except it’s not that simple. A growing body of academic and industry research has cast doubt on whether these “intermediate tokens” are a faithful representation of an LRM’s inner workings. Instead of being auditable receipts or accurate reports, they can appear more like what the Arizona State University researcher Subbarao Kambhampati calls “mumblings” — bits of language, yes, but ones whose meaning may be entirely incidental to any reasoning that might have occurred. Kambhampati’s lab showed in 2025 that fully replacing a model’s correct “traces” with incorrect or irrelevant ones didn’t degrade its performance on a formal reasoning task. Meanwhile, training the model only on correct trace data still led it to occasionally generate invalid records of its reasoning — even when it produced a correct solution to the original problem it was given. A 2024 paper from researchers at New York University showed that “meaningless filler tokens” — literally, strings of dots — could function effectively in place of a human-readable “chain of thought.”
William Merrill, one of the authors on that paper and currently a professor at the Toyota Technological Institute at Chicago, put the matter plainly: “There’s no guarantee the chain of thought has to be meaningful in any sense.” Pavel Izmailov, a researcher at NYU who also works for Anthropic (and was part of its original reasoning-model team), said he doubts that reinforcement learning — a typical training method for LRMs — even incentivizes models to produce faithful chains of thought in the first place. “I mean, maybe it will,” he told me. “But I would say the chances are not very high.”
OK, so the linguistic content of reasoning traces may be dubious. But surely the tokens themselves must play a role in producing the model’s outputs? (Think of a pinball machine: It runs on coins, not the words “In God We Trust.”)
Not so fast. A 2025 paper from Northeastern University and the University of California, Berkeley on frontier open-source LRMs showed that between 30% and 60% of their “thinking steps” had “minimal causal impact” on the answers the models produced to benchmark math questions. Chop half of them out, and a model’s performance barely suffers. “We want to be careful when we review these chain-of-thought prompts because they may not be linked to the final output,” said Weiyan Shi, one of the study’s authors.
So reasoning traces, the very things that supposedly distinguish LRMs from the mere next-word-predicting LLMs, are not necessarily either meaningful or causal to a model’s … reasoning? I’m no philosopher, but this seems to stretch the meaning of “reasoning” beyond its tensile strength. Kambhampati’s research group sounded frankly fed up in the title of their position paper on the subject (presented at the 2026 International Conference on Machine Learning, one of the field’s most prestigious academic gatherings): “Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!”
To be clear, Kambhampati, a former president of the Association for the Advancement of Artificial Intelligence, with a background in AI planning algorithms, doesn’t deny that LRMs can work (when they work). “We are in wondrous times,” he told me, when I asked what he thought of OpenAI’s 2026 victory in solving the famous unit distance problem in math. If he has a bone to pick, it’s with what he sees as a rush in both academia and industry to embrace overly convenient explanations.
“Many ideas that have been proposed [about] the sources of strength [of these models] have been misunderstood or mischaracterized,” he said. “There’s this general mindset that says, ‘Let’s go ahead and claim certain abilities, because eventually that might become true anyway.’ And my sense is: That’s not science. That is investment.”
On the other side of the AI-reasoning fence, the disdain seems to be mutual. “These ‘scientific’ papers from last summer — I would put this in big, big air quotes,” said Sébastien Bubeck, a member of OpenAI’s technical staff (and a prominent evangelist for the company’s reasoning models among scientists and mathematicians). He called earlier Apple results critiquing AI reasoning “wrong,” claiming that they were due to a training quirk in models that are now obsolete. “Modern models starting with GPT-5.5 do not suffer from this issue,” he said. “It would be interesting to revisit those results.” (Apple did not make its researchers available for interviews.)
![]()
Here’s the thing: Nobody denies that AI reasoning models can, indeed, produce significant and accurate results. Furthermore, every researcher I spoke to acknowledged that negative findings about the models’ capabilities on certain reasoning tasks (especially those of smaller, open-source LRMs) may not always generalize to the latest-and-greatest AI products. Their inner workings remain trade secrets. But if we’re disinclined (as I am) to simply dismiss contradictory evidence about the mechanisms driving AI reasoning, the question remains: How do we account for it?
Kambhampati, as it turns out, is interested in doing exactly that. “I’m not negative. I just sound negative because everybody else is way too positive,” he said. “In science, you have to actually understand what the current thing does and what it cannot do.”
One straightforward reason state-of-the-art LRMs work, he told me (a point also echoed by Mitchell), is that they’re often surrounded by “normal” software that guides and verifies their outputs. Agentic AI systems, which have transformed software engineering since the fall of 2025, work this way. So does Google DeepMind’s AlphaProof Nexus, which relies on Lean, an automated theorem-proving tool. But Kambhampati is more interested in making sense of stand-alone reasoning models that rely solely on their self-generated reasoning traces — “the ‘think’ part,” he said.
The “think” part is what OpenAI, for one, is doubling down on. When I asked Bubeck if the splashy unit distance proof was produced with methods outside the LRM’s own chain of thought — perhaps with Lean verifying its results — he seemed to find the question almost nonsensical.
“It’s not like we’re making a mystery of it,” he said. “We have released the chain of thought. You can just go and look at it. The whole point is that the model is reasoning like a human would. And when humans reason, we don’t use Lean.” Technically, OpenAI released a “rewritten summary” of the model’s chain of thought produced by two human experts using Codex, another OpenAI model. Since 2024, the company has not publicly revealed “raw” chains of thought from its reasoning models, a policy also adopted by Google DeepMind and Anthropic.
Kambhampati’s analysis begins in a surprisingly similar place: with the idea that LRMs are just LLMs with more specific training. “There is no extra magic,” he said. But he diverges sharply from there. “It doesn’t make sense to me that an LLM would actually do a step-by-step description of what it is [reasoning] before giving the solution — because that’s a much harder task than just guessing the solution, given the way that LLMs are trained.”
His working hypothesis is that an LRM, like its LLM precursors, performs what he calls “approximate retrieval” across its vast training corpus: “somewhere in the middle” between pattern matching and reasoning, he said, but closer to the former. The role of “thinking tokens,” then, isn’t to narrate an actual chain of thought (because there isn’t one). Instead, it’s to load up the model’s context window in a way that makes it more likely to predict, or “approximately retrieve,” reasoning-shaped strings of text.
Kambhampati compared this process to mumbling words to yourself to jog your memory: It barely matters what the words are (though related ones may help), as long as they knock loose something useful. An LRM’s vast “memory” includes all the call-and-response-like examples of written reasoning it was trained on, mulched into numerical “embeddings” that encode their similarities and differences (plus other inscrutable associations) as geometric relationships in a high-dimensional space. Probabilistically arriving at an answer within that space may involve intermediate tokens whose embeddings map to coherent-looking “thoughts” in plain English, but not necessarily. They could be bits of other languages. They could be fake exclamations like “aha.” Under the right conditions, they could just be dots.
“Whether the [embedding] actually corresponds to a single word or not” — much less a faithful reasoning process — “is beside the point,” Kambhampati said.
This framing could help explain both the odd “BS”-ness of some chains of thought and the fact that they can elicit accurate outputs anyway. It would also neatly account for LRMs’ steady improvement in coding and math — what AI researchers call “verifiable domains.” Code runs, or it doesn’t; proofs are either correct or not. These binary conditions and the written steps associated with them can create convenient training signals for LRMs. The model doesn’t have to learn or reliably apply a general reasoning process, Kambhampati said; it just has to absorb enough examples of what the steps look like to predictively mimic them on its way to “stitching together” a plausible result that can then be verified.
The limit of a reasoning model’s training and step-following capability, known as the “inference horizon,” Kambhampati added, was what Apple researchers exposed with their “Illusion of Thinking” paper in 2025. Newer models have appeared to push this horizon further, albeit jaggedly. “Most of the time they probably are not learning the algorithm” associated with a reasoning process, he said. It’s much likelier that they are leveraging an ever-enlarging set of examples and clever reward signals.
Kambhampati hardly considers his case closed, and neither do I. But it’s a start — and one I find plausible, given that other researchers have also used similar “it’s the training, stupid” approaches to demystify AI behavior. Still, there was an elephant left in the room: How much does it matter whether or not we can accurately observe, characterize, and validate the processes at work inside large reasoning models?
The honest answer, according to Mitchell, is that it depends. “Think of AlphaFold,” she said, referring to Google’s AI tool for predicting protein structures. “It’s doing some kind of incredibly complex statistical associations. We don’t know what they are, but they seem to work. These things are [already] black boxes, even without a ‘reasoning trace.’” If LRMs can supercharge mathematics research the way AlphaFold did for computational biology, this line of thinking goes, why not embrace them, idiosyncrasies and all, and just verify the results? “My perspective is: We’re trying to be useful. We’re trying to build these models so that they can solve problems that matter, so that we actually accelerate scientific research,” said Bubeck. “It’s more interesting and more productive to talk about what they can do, rather than, ‘Oh, but they can only do that because of X [reasons].’”
But as Mitchell also points out, the possibility that an LRM could be “right for the wrong reasons” has an obvious relevance to the future of doing research. “You want the right answer for the right reason, so you can trust these things,” she said, and not just in verifiable domains.
Tal Linzen, a researcher at NYU and Google whose Computation and Psycholinguistics Lab published results similar to Apple’s “Illusion of Thinking” paper, said that “you want an AI system to be able to apply an algorithm reliably, regardless of whether you call [it] reasoning or not.” Treating chains of thought too reverently — even when their results are verifiable — could also prevent scientists from discovering even better ways of biasing LRMs toward accurate outputs. “We may be leaving some opportunities unexplored,” said Pradeep Dasigi, a researcher who helped train open LRMs at the Allen Institute for Artificial Intelligence. Kambhampati, unsurprisingly, puts it in even starker terms: Taking the meaning of AI reasoning traces seriously, he said, was a scientific “rabbit hole,” akin to believing in geocentrism or the ether.
Harsh, perhaps, but he has a point. Those incorrect mental models made intuitive sense at the time, just as chains of thought do now. When an LRM produces a correct answer — along with pages of “thoughts” showing how it got the result — intuition tells us that the two must be linked. It’s hard to imagine that process and outcome may have little to do with each other. But in the 1990s (in an episode Mitchell and Izmailov both brought up), it was hard to imagine how brute-force search could beat world champ Garry Kasparov at chess. And in 2023, it was hard to intuit how a giant pile of matrix multiplications could write in iambic pentameter. For most of us, these just weren’t thinkable thoughts. Until, suddenly, they were.
![]()
In summer 2024, just months before the first LRM appeared, Mitchell turned me on to a concept that I keep returning to in my AI reporting: “wishful mnemonics.” The phrase was first used all the way back in 1976 by the computer scientist Drew McDermott, in a paper with the epically grouchy title “Artificial Intelligence Meets Natural Stupidity.” I’ll quote the same passage Mitchell did:
A major source of simple-mindedness in AI programs is the use of mnemonics like “UNDERSTAND” or “GOAL” to refer to programs and data structures. … If a researcher … calls the main loop of his program “UNDERSTAND,” he is (until proven innocent) merely begging the question. He may mislead a lot of people, most prominently himself. … What he should do instead is refer to this main loop as “G0034,” and see if he can convince himself or anyone else that G0034 implements some part of understanding. … Many instructive examples of wishful mnemonics by AI researchers come to mind once you see the point.
This is how I make sense of AI reasoning. LRMs, chains of thought, thinking tokens: It’s wishful mnemonics all the way down — a heady mix of shorthand and suspended disbelief, like Oprah-style “manifesting” with a computer science spin. This isn’t necessarily a dig; all novel research likely requires some version of this mindset just to get off the ground. It certainly doesn’t mean AI reasoning can’t or doesn’t work. But the “wishful” part seems to be as powerful as ever.
“We react to language in a way that is very anthropomorphizing. That’s just the way that we humans work,” Mitchell told me. Much of the contentious research activity around AI reasoning, she said, “is par for the course. But in other ways, there’s a lot of very unscientific aspects to it.” Or, as Kambhampati put it, “A fake theory is worse than admitting that we don’t have a theory.”
In any case, we have to call it something while we figure out what it is. I don’t foresee always reaching for the air quotes around AI reasoning, any more than I’d put them around the “horse” in horsepower. LRMs are like engines: They require fuel, emit exhaust, and go fast. Still, when I describe the oomph my Toyota can deliver when I step on the gas, it’s not because I believe there are little hooves pounding away under the hood. Until a clearer scientific account emerges of what’s going on under the hood of AI reasoning models, I’ll regard their horsepower in a similar spirit — even as the engines roar.