Reasoning Is Not Language
On fluency, the license to decide, and where models belong
19 JULY 2026 · ABOUT 64 MINUTES
Contents
I. The wager
In a theoretical sense, a clinician in the midst of raucous rounds could use a frontier AI model to produce a recommendation, scrawled in text across a screen, that describes in detail how a certain patient situation should be handled — which, again in a theoretical sense, could be better sourced, more current, and more complete than anything she could arrive at de novo.
The recommendation would be delivered, as we have all become accustomed to, in the appropriate register and characterized by a level of self-certitude perfectly calibrated to instil confidence in the clinician — a confidence indifferent, as it happens, to whether the recommendation is correct, or whether some necessary nuance has been omitted.
However, no matter how florid the prose, or how self-certain the claims, this is an unworkable system as instantiated now, and as it could ever be (in our view) instantiated in the future. Because neither the model nor, more importantly, its creator can be sued for an improper recommendation, or dismissed for it, or made to answer, perhaps years later, for it. Whereas a clinician can be, and often will be, asked to justify her actions — whether as an administrative aside or as part of an investigation into misconduct. And saying the machine said so with no substantive rationale behind it will never be an acceptable answer.
What would count as an acceptable answer is one derived from a clinical audit trail: a record of how the recommendation was arrived at, step by deterministic step, all from logic rooted in an agreed-upon clinical standard — in other words, logic that runs in such lockstep with the standard that, between the two, there is no room left for variance unless the clinician, exercising her judgment, determines that variance is warranted and documents as such.
But for every frontier model now in service, it is language, not logic, at every surface one can inspect. It is in language that these systems are asked to ‘reason’, in language that their ‘reasoning’ is made visible, and in language that the result is handed to whoever asked for it.
Whether the fluent paragraph output by a frontier model onto a clinician’s screen is a medical judgment, or a simulacrum of one, comes down to a question about language: whether a model built entirely of language can be made to reason in it, reliably enough that the avoidable error does not reach the patient.
This, as it happens, is a mere reformulation of one of the oldest questions in philosophy that now carries a more practical weight in the age of AI — because if we think in language, the model’s fluency may be the reasoning itself. But if reason stands apart from language, then a probabilistic model will perhaps always need muzzling before it is let near a patient.
Aristotle thought the linkage of language with thought intolerable. In De Interpretatione he distinguished spoken symbols, which differ across peoples, from the mental affections those symbols signify, which do not1 — a distinction that presupposes the very separation now at issue, namely that there is something a thought is, prior to and independent of the words in which we happen to report it.
But this has not been treated as a fait accompli — rather, this denial has been contested ever since. Whorf held that the structure of a language conditions the very thought of its speakers; Carruthers has made natural language the medium of domain-general reasoning; Lupyan has put it at the center of cognition rather than at its periphery.2
However, athwart these counters stands the tradition of Fodor and his acolytes, whose objection was as much logical as empirical: thought must be conducted in some compositional symbol system, but that system cannot be any natural language, since one would have to possess it already in order to learn English.3 Mentalese, if it exists, is not English written small.
For most of the long history of this argument it stayed where such low-stakes arguments typically stay: in the province of philosophy, conducted by disputation and introspection and the occasional thought experiment, its stakes low because everyone concerned agreed that no experiment would ever settle it.
Two things have since reinvigorated this centuries-old stalemate. The first is that neuroscience can now localize the language faculty in a living brain and ask it directly whether it takes any part in inference. The second is that the question has acquired a new level of practicality since which answer a frontier lab believes in now dictates what it builds, so that the frontier labs persuaded that reasoning lives in language will scale a large language model and wait for the proof of their answer to reveal itself.
To underline this, the proposition that a model trained to predict text will, at sufficient scale, acquire the capacity to reason is neither a neutral technical hypothesis nor without consequence. Rather, it is a commitment — rarely acknowledged as such by those who hold it — that reasoning is, or emerges from, linguistic competence. Thus, in effect, most frontier labs have placed an extraordinary wager on one side of a very old argument (whether they realize it or not).
In our view, it is a losing wager, though not for the reason one might expect. The empirical question has not gone against these labs — that question may never be settled, and we will show at length how tumultuous the debate remains.
However, the wager loses in a narrower, more practical, and more stubborn sense: in the domain of our interest, medicine, it collides with how the work has always had to be done, and always will be done, and to suppose a model could ever reason sufficiently to drive clinical decision-making is to chase a chimera (at least, as we’ll discuss, in the sense most labs and those that provide wrappers around them imagine).
Our view, despite the academic back-and-forth to come, is that the inference which reaches a clinical decision must be carried on a substrate that is not natural language — and, further, that this does not depend on the models being any less capable than their most enthusiastic advocates claim (which is fortunate, because they are not).
As we’ll discuss in the sections to come, there are two arguments here that run in parallel. The first — which occupies the next five sections, and which we happen to believe — holds that a model has no business reaching a clinical determination because it cannot reliably reason its way to one. That argument depends on the models remaining roughly what they are (not where they are), and we would not want to bet a company on that fact — in §VI we set out, at some length, why a reasonable person might conclude this argument is already offsides.
The second argument does not depend on either side of the argument that preceded it. It asks what a clinician may rely upon in the moment she signs her name — what the, for lack of a better word, certificate beneath the signature would have to be — and the answer turns on nothing that any lab can change. A model good enough to sink the first argument is the model that paradoxically makes the second one more urgent (and is a not so hidden driver behind the statutory language increasingly being implemented in North America).
The claim, stated once and plainly, is a double asymmetry. The better artificial intelligence becomes, the more medicine requires that its own standards be executed deterministically — and, in turn, the better (more comprehensive) that execution becomes. The first half of that sentence will strike many readers as perverse, and it is the burden of §VIII to prove it. The second half, likewise, will strike many readers as a concession to technology they did not expect from an argument of this shape, and it is the burden of §X to explain it. Taken together what’s described herein is an architecture in which one may be as bullish about the frontier as one likes but arrive at the same outcome — which is not a retreat from artificial intelligence but an embrace of it, through the use of an appropriate harness, to maximize clinical impact.
Part One
Where Reasoning Happens
Reasoning and language are separate systems, and fluency is therefore no evidence of inference — this is our more academic argument we hold and press while conceding that it is the weaker of the two arguments we will make since it depends on the models not making a fundamental step-change in capability. The argument of Part Two does not depend on it at all.
II. The silent network
To begin, the most persuasive evidence, on our side of the ledger, comes from Evelina Fedorenko and her collaborators at MIT, who have spent two decades establishing what the human language network is and, more usefully, what it is not.
The network, to use their frame, is characterized as a set of left-lateralized frontotemporal regions, identifiable in the individual subject, robustly engaged by lexical and syntactic processing across languages, modalities, and tasks.4
What has become steadily more pronounced in their recent work is the network’s selectivity. These aforementioned ‘regions’ are largely silent during arithmetic and algebra. They are silent during the comprehension of computer code. They are silent during music perception, working memory, and social reasoning.5
This means demanding non-linguistic cognition is carried elsewhere — substantially by the domain-general Multiple Demand network, a distinct frontoparietal system whose damage reduces fluid intelligence and whose activity scales with task difficulty rather than with linguistic content.6
Of note is that formal logical reasoning had, until recently, gone suspiciously unexamined. The gap was odd, since logic is the one domain where the language-as-medium hypothesis has always looked most secure: syllogisms are made of sentences, and their validity seems to turn on their structure much as a sentence’s meaning turns on its syntax.
Kean, Fedorenko and colleagues closed the gap this year, in PNAS.7 In healthy adults, they showed that the language network shows no reliable response to deductive syllogistic reasoning — the contrast between the harder Modus Tollens and the easier Modus Ponens, which isolates deductive load while holding linguistic content constant — nor to non-verbal matrix reasoning, despite pronounced behavioral difficulty effects in both.
Inductive reasoning does elicit a small response at the network level, but it is four times weaker than the language contrast, and the critical condition produces activity at or below what the same regions show while the subject reads a list of nonsense words. The language system, in other words, is not doing the logic — and in the case of deduction, it is not doing anything at all.
The complementary evidence is all the more striking, because it is evidence by subtraction. Two patients with profound aphasia — one following left-hemisphere stroke, one following subdural empyema and meningitis, both with extensive peri-Sylvian damage, both classified as severely agrammatic, both performing at or near chance on multiple assessments of grammatical comprehension in speech and in writing, both unable to determine who splashed whom in a reversible sentence — were given the same reasoning tasks. They solved them. One solved thirty-nine of the forty induction rules presented to him; the other, nineteen of twenty-five. On the standard matrix-reasoning instrument, both scored well above the age-matched population mean, with T-scores of seventy-three and sixty-eight against a normative mean of fifty.8
What was demonstrated, in brief, was a double dissociation: the language network active for language and inert for logic; the aphasic brain destroyed for language and intact for logic. Correlational findings in this literature — the observation that aphasic patients often perform poorly on reasoning tasks — have long been used to argue the contrary case, but correlations of that kind are cheaply produced by anatomical proximity, since a stroke does not respect functional boundaries. Dissociations are not so burdened.9
SOURCE · Kean, Fung, Jaggers et al., PNAS 123(28): e2520095123 (2026). Effect sizes and p-values at Table 1; patient scores at Results and Fig. 2D; WASI-II norms, T-mean 50, SD 10. n = 2 patients — the study’s first stated limitation.
inspect
| contrast | effect size | p |
|---|---|---|
| Language | 1.55 | — |
| Induction | 0.26 | < .05 |
| Matrix reasoning | 0.15 | .12 |
| Deduction | −0.04 | .71 |
| patient | raw | T-score | vs norm |
|---|---|---|---|
| S.A. | 25/30 | 73 | +2.3 SD |
| G.S. | 26/30 | 68 | +1.8 SD |
Effect sizes from Table 1; patient scores from Results and Fig. 2D; WASI-II matrix reasoning, normative mean T = 50, SD = 10.
Of course, objections exist, and the fairest catalog of them is Varley’s own since she has worked with such patients for a quarter of a century and is among the authors of the study just described.10
There’s no getting around that the samples here are small (two patients, in this case). A man with aphasia this severe may fail a task because he has not grasped what it asks of him rather than because he cannot do it. Likewise, the impairment is acquired, which means these men had language before their strokes, so that a mind which has lost language cannot testify to what language did while that mind was being assembled.
And there is the Chomskian objection, that a lesion disturbs the use of a capacity and not its possession — to which Varley’s reply, borrowed from Whitaker, is that a competence no brain injury can reach is not a property of the brain at all. We think the first objection and the third bite on the subtraction evidence and leave the imaging where it stands, since the imaging needs no patients: in the intact brain, with the grammar unimpaired, the language network still sits out the syllogism. The second we accept, since it marks the boundary of what this evidence can be asked to do, which is to describe how a built mind runs rather than how it came to be built.
With all that said, this dissociation does not stand alone: it is the latest in a convergent series establishing that mathematical reasoning, causal reasoning, theory of mind, intuitive physics and code comprehension all proceed without recruiting the language system. A companion investigation from the same group suggests that deduction recruits machinery distinct from both language and the general-purpose executive network — that formal reasoning, in other words, has an organ of its own.11
Fedorenko, Piantadosi and Gibson have drawn the threads together in the obvious way: language is optimized as a system for communication — an efficient code for transmitting thoughts between minds — rather than as the medium in which thoughts are computed.12 It interacts constantly with the reasoning systems, since one must have something to say. But the interaction occurs at an interface. The inference happens somewhere else.
III. The format of thought
So, if not language, then what carries thought? The leading answer is the Probabilistic Language of Thought, on which mental representations are compositional symbolic programs over concepts, and reasoning is inference over those programs.13
The connection to the evidence of §II above isn’t incidental: the rule-induction task on which the aphasic patients performed so well is itself a program-induction task. The participant sees that [5, 7] becomes [7, 5, 7] and must infer the transformation — add two to every element, retain only the unique elements, rotate the list by three — each rule expressible as a short program.14
The leading account of the human reasoning substrate, in short, is that it is compiled, compositional and program-like — and distinct from the linguistic surface on which its results are eventually reported. Which all tells us what a machine that reasons reliably ought to be doing: constructing something program-like, checking it, and executing it, rather than generating a plausible linguistic continuation and hoping valid inference is somewhere inside (we’ll leave the similarities of all of this to Compiled Health’s approach as an exercise for the reader).
Ultimately, the opposing tradition should not be dismissed out of hand. Language plausibly does real cognitive work: it compresses culturally useful concepts into transmissible units, it scaffolds the acquisition of abstract and relational structure, it lets one offload intermediate results onto external symbols. But that language helps one learn a concept is a wholly different claim from the proposition that one thinks in language — just as learning a concept from photographs does not thereafter render one’s thinking photographic. The facilitation thesis survives the evidence, but the constitution thesis does not.
So, what has been shown is that reasoning can proceed on a substrate that is not natural language. It has not been shown that it must, and it has certainly not been shown that a machine cannot reason in language.
After all, the brain is an existence proof, not an impossibility proof. That a biological system evolved to route inference around its language faculty tells us that the two are separable. However, it tells us nothing whatsoever about what an artificial system, built differently and trained differently, may or may not achieve inside the linguistic medium.
To move from the first claim to the second is to commit the oldest error in the philosophy of mind, which is to mistake the way a thing happens to have been made for the only way it could be made.
Neuroscience does not prove that machines cannot think, or could never be built to. It proves, by our lights, something smaller and far more useful: that thinking and speaking are different jobs.
IV. Well-formed and wrong
The bridge from neuroscience to large language models was built by the same group. Mahowald, Ivanova, Blank, Kanwisher, Tenenbaum and Fedorenko distinguish formal linguistic competence — command of the rules and statistical regularities of a language — from functional linguistic competence, the capacity to deploy language in the world, which necessarily recruits reasoning, world knowledge, situation modeling, and social cognition.15
Their assessment, in 2024, was that large language models had substantially mastered the first while remaining uneven at the second, which is unsurprising: deploying language in the world requires machinery a language model does not (at least at time zero) have.
Our account predicts the empirical literature: systems all but flawless at producing well-formed, fluent, plausible prose fail — in a characteristic and diagnostic way — at the inference that fluent prose is supposed to carry. The signature of the failure is sensitivity to variation that is irrelevant to a reasoner and highly relevant to a pattern-matcher.
The cleanest academic demonstration is GSM-Symbolic, which regenerates grade-school arithmetic problems from symbolic templates. Change nothing but the numbers and accuracy falls across every state-of-the-art model tested. Then go a step further and insert a clause that is topically germane but logically inert — the ‘NoOp’ condition — and accuracy falls still further.
SOURCE · Mirzadeh, Alizadeh, Shahrokhi et al. (Apple), GSM-Symbolic, ICLR 2025 (arXiv 2410.05229), Table 1. CONTESTED · a 2026 re-analysis (Długosz et al., arXiv 2605.28700) argues the drops shrink toward the noise floor once distractors are audited for true irrelevance; the essay treats the finding as suggestive, not settled.
inspect
| model | Symbolic | +1 clause | +2 clauses | + no-op | drop |
|---|---|---|---|---|---|
| GPT-4o | 94.9 | 93.9 | 88.0 | 63.1 | −32 |
| o1-preview | 92.7 | 95.4 | 94.0 | 77.4 | −15 |
| Phi-3.5-mini | 82.1 | 64.8 | 44.8 | 22.4 | −60 |
| Gemma2-9B | 79.1 | 44.0 | 41.8 | 22.3 | −57 |
Accuracy (%), GSM-Symbolic Table 1. The clause conditions add material to the problem text; the arithmetic required is unchanged.
In other words, even though the problems stay trivial, and a child who understood them would not be thrown, the models are thrown, and (sometimes) thrown hard — which licenses the authors’ natural conclusion that what runs in these systems is not formal reasoning but the recall of reasoning-shaped patterns.16
The same signature shows up elsewhere. Transformers collapse multi-step compositional problems into shortcut heuristics and degrade sharply as the required depth mounts; behavior tracks the statistics of the next-token objective rather than the logic of the task, so that a problem phrased against the grain of that objective can break a system that solves harder problems phrased with it (yes, progress has been made here; no, the fundamentals have not changed).17
To be clear, these are claims about mechanisms, not about whether the outputs are ever right. Often they are right — but, and this must be underscored, that is where the trouble starts because any system (LLM or otherwise) that is usually right for reasons you cannot inspect or duplicate is far more dangerous than one that is reliably wrong.
Not everyone reads this record the same way. The strongest of the fragility papers has drawn substantive rebuttal, and a re-analysis of GSM-Symbolic argues that some of the apparent degradation sits within statistical noise.18 Although the perturbation sensitivity appears real and has been widely replicated, no single striking result can ever be decisive. Both can be true, and the reader who wants the strongest case against our overall view will find it in §VI not here.
V. Propose, then prove
If language is the wrong place to carry the inference, the corrective is to keep language at the interface and route the inference through an engine built for it (or design around it in toto but we’ll leave that aside). Two literatures have been quietly doing exactly that for years well before the current wave made it fashionable.
The first is neuro-symbolic. A cluster of systems — Logic-LM, SatLM, LINC — takes a problem stated in natural language, parses it into a formal specification (first-order logic, a satisfiability instance, a constraint program), dispatches it to a symbolic solver, and renders the solver’s output back into prose. The reported gains over language-only chain-of-thought on formal-reasoning benchmarks are real where the problem admits a formal specification, for a reason that is not mysterious: the deductive step is performed by a sound engine rather than approximated by token prediction. Related work pushes toward certification, guiding the model to emit reasoning that an external checker can verify step by step.19
The second is program synthesis and policy extraction, and its logic is more instructive: a high-capacity, often stochastic, search process discovers behavior, and the discovered behavior is then retained as an explicit, checkable artifact. This approach has been implemented across a decade worth of systems — DreamCoder growing libraries of symbolic programs under neural guidance, FunSearch pairing a model proposer with a deterministic evaluator, AlphaDev discovering sorting routines that then ship, as ordinary code, in standard libraries used by people who will never meet the agent that found them, PIRL and VIPER extracting formally analysable policies from neural controllers.20
The governing ideal is Necula’s proof-carrying code: one accepts an artifact not because its generator is trusted, but because a small checker can verify the evidence traveling alongside it. Rudin’s argument for interpretable models in high-stakes decisions is the same conviction stated as simple policy — when the decision matters, prefer a model whose reasoning can be inspected to a black box explained after the fact.21
Reduced to its stages: a generator proposes; a selector filters artifact candidates against evidence and specification; retention fixes the survivor as a versioned artifact; verification checks it. What the end user (e.g., a clinician) is given to act upon is thus the verified artifact, not the generator — a system exposed to her as:
where monitoring in deployment feeds evidence back to the generator and the selector, but any revision must re-enter through retention and verification before it can reach an end user. The discovery may be as wild and probabilistic as you please; what the end user is shown has been checked.
The showpiece here, as is often the case in early proof of concepts, is olympiad mathematics. AlphaGeometry solved twenty-five of thirty olympiad geometry problems by pairing a neural language model — which proposes the creative auxiliary constructions that had defeated automated geometry for decades — with a symbolic deduction engine that guarantees the validity of every step.
AlphaProof generalized the pattern, operating inside the Lean theorem prover so that each step of each proof is machine-checked; together the two systems reached silver-medal standard at the 2024 International Mathematical Olympiad.22 These are, by design, systems in which the neural model supplies intuition and the symbolic solver supplies certainty with neither asked to do the other’s work.
inspect
| year | system | generator proposes | verifier keeps |
|---|---|---|---|
| 2021 | DreamCoder | neural program search | task tests + library |
| 2022 | AlphaTensor | RL tensor algorithms | exact equivalence |
| 2023 | AlphaDev | RL assembly programs | correctness tests |
| 2023 | FunSearch | LLM program candidates | scoring evaluator |
| 2024 | AlphaGeometry | LM constructions | symbolic proof engine |
| 2024 | AlphaProof | LM tactic candidates | Lean type-checker |
| 2025 | Deep Think (IMO gold) | LM natural-language proofs | — none external; see §VI |
Venues: Ellis et al. 2021 (PLDI); Fawzi et al. 2022 (Nature); Mankowitz et al. 2023 (Nature); Romera-Paredes et al. 2024 (Nature); Trinh et al. 2024 (Nature); Google DeepMind 2024, 2025. Citations in §V–§VI.
SOURCE · systems and pairings as published by their labs; dates are first public results. The one open point is the exception the essay concedes, not a member of the pattern.
In general, evidence comes not from the critics of the frontier labs but from the labs themselves — the Coconut architecture proceeds from the observation that the language space may not be the optimal space in which to reason (e.g., that most word tokens exist to secure textual coherence rather than to carry inference) and modifies a language model to reason in a continuous latent space, feeding hidden states forward as inputs rather than decoding them into words.
On logical-reasoning tasks requiring search, latent reasoning outperforms verbalized chain-of-thought while consuming fewer tokens, because the model can hold several branches open rather than committing prematurely to a single articulated path.23 The authors’ framing of their own contribution is largely our house view, arrived at independently and from the opposite direction: let the model reason free of linguistic constraint, and then translate its findings into language.
Now, while Coconut proposed that a language model ought to reason outside language, in July of 2026 Anthropic reported that, left to itself, it already does. Using a new interpretability technique — the Jacobian lens — the company identified what it calls a J-space: a small, privileged region of internal activity, no more than about a tenth of the model’s activation variance in any layer, where concepts are held, manipulated and reasoned with before any of them reach the output.24
The finding that matters here is the ablation: remove the contents of the J-space, and the model continues to speak fluent, well-formed, grammatical English while its multi-step reasoning degrades sharply. Fluency runs in the basement, like breathing; deliberation runs in the workspace. They are separable, and when they are separated it is the reasoning that goes.
This is the dissociation of §II, found inside the model. It was not deliberately designed; it emerged in training and was framed by its discoverers through the lens of the very neuroscience we have been describing whose principal architects were invited to comment. With the implication being that a language model, left to itself and given enough pressure to reason, appears to arrive at a division of labor that resembles the one evolution arrived at: a large automatic apparatus for producing language, and a small privileged space, distinct from it, where the deliberation is done.
We would not rest a case on a single interpretability result, and its authors are careful about what the technique does and does not show. But the resemblance is difficult to unsee. And it is, as we are about to discover, the most dangerous piece of evidence to-date.
VI. The case against us
Now, we’ll trot out the other side of the argument — though, as we said at the outset, our priors notwithstanding, we remain agnostic about the whole debate as a practical matter (because no matter how robust the J-space, or whatever supersedes it, it still cannot be used alone at the bedside — not now, not ever).
We’ll begin with reinforcement learning — a term that has been so used and abused it has lost most of its meaning. DeepSeek-R1 established — and Nature has now published — that reasoning behavior, including self-verification, backtracking and dynamic strategy revision, can be incentivized in a language model by reinforcement learning against verifiable rewards, without supervised reasoning traces of any kind.25 We thank, since DeepSeek won’t, DeepMind for their contributions on this.
The gains were real and they held up under scrutiny, which makes the moral an inconvenient one: take a single undifferentiated language model, with no symbolic organ anywhere in it, train it against outcomes that can be checked, and its language-medium reasoning grows dramatically more reliable — though not unimpeachable, which is rather the point for us.
As a reminder, in 2024, the neuro-symbolic architecture described above — AlphaProof and AlphaGeometry, formal specification, machine-checked proof, etc. — reached silver at the International Mathematical Olympiad. However, in 2025, Gemini Deep Think reached gold, solving five of six problems, using end-to-end natural-language reasoning, with no translation into a formal language and no external symbolic engine anywhere in the loop.26
On the very benchmark that had motivated the formal architecture, the language-medium system beat it. We concede that without qualification, because it is the single strongest piece of evidence in the whole literature against the position we fleshed out in the prior section.
We would be remiss, though, not to note the conditions under which the result was obtained. A gold medal at the Olympiad is won across two sessions of four and a half hours, against six problems, by a system permitted to deliberate for as long as the rules allow and to consume whatever compute its sponsor is willing to fund. That is a magnificent achievement. It is also a set of conditions that bears no resemblance whatever to a Tuesday morning clinic and runs roughshod over the principle that how an output is arrived at is as important as the output itself.
Nor can a retreat be made to the claim that reinforcement learning merely sharpens what was always latent. It is tempting to make it, because a much-discussed result — presented at NeurIPS by Yue et al. — appears to establish exactly that: measured by pass@k at large k, reinforcement-trained models beat their base models on a single attempt but are overtaken by those base models when both are permitted many attempts with the reasoning paths of the trained model turning out to lie within the base model’s sampling distribution.27 Stated in that metric, the finding is a crossover:
SCHEMATIC · The crossover is the qualitative finding of Yue et al. (2025) on RL’s reasoning boundary; the curves are illustrative, not fitted to any single benchmark. Contested directly by ProRL (Liu et al., 2025), which reports genuine expansion. See §VI.
inspect
| attempts k | RL-tuned % | base % | leader |
|---|---|---|---|
| 1 | 74.8 | 31.7 | RL-tuned |
| 4 | 88.0 | 68.0 | RL-tuned |
| 16 | 88.0 | 87.6 | RL-tuned |
| 32 | 88.0 | 90.9 | base |
| 1024 | 88.0 | 99.0 | base |
Generating functions: RL(k) = 88·(1 − 0.15k+1); base(k) = 99·(1 − 0.68k+1), k as powers of two. Schematic — the shapes carry the Yue et al. finding; no benchmark is being reproduced.
On this account reinforcement learning teaches a model to sample better, not to reason newly, and the reasoning boundary in fact narrows as training proceeds.
The temptation to double down on what’s always been latent should be resisted, and those who yield to it watched closely, because the result is contested. ProRL, from NVIDIA and presented at the same conference, argues that prolonged reinforcement learning — with divergence control, reference-policy resetting, and a sufficiently diverse task suite — uncovers reasoning strategies genuinely beyond the base model even under heavy sampling, including on tasks it fails outright no matter how many attempts it is given.28
Both groups have their own independent motivations (as it happens, the financial incentives of both align with their findings). Nevertheless, here we have two papers, at one (renamed) venue, the same metric, and opposite conclusions. The honest position is that whether reinforcement learning expands machine reasoning or merely concentrates it is an open question, and an argument which requires this to resolve in one particular direction amounts to an argument that rests on a precarious precipice.
But there is a more foundational issue at play: if a language model spontaneously develops a privileged internal workspace in which reasoning is conducted apart from language production — and if that workspace can now be read and even intervened upon — then the sharp architectural dichotomy discussed here begins to look less like a design squabble and more like a distinction the models are quietly dissolving on their own (in other words, a distinction that soon will be without a difference).
So, if we continue this chain of thought, pun intended, why insist on an external symbolic engine (or something analogous in function to one) if the model could grow an internal one? And why insist that its reasoning is unauditable, if the reasoning can be inspected directly, beneath the level of the words?
The objection is a fair one on its face, if speculative about future model advancements, but the answer turns on a conflation buried inside it: between being able to see what a system was thinking and being able to rely on what it produced every time. Those are different properties, and only one of them is any use to someone who must answer for a decision — all the more so when the clock is running. Indeed, even though the above academic cut and thrust is commercially less important to us, our views on it are informed by us working backwards from a set of simple pre-requisites (reliability and verifiability) and constraints (time and budget).
Part Two
What a Decision Requires
The question now changes from what a model can do to what a license is, and the answer bears no relation to model capability: a license is a standing liability, a model has nothing it can be made to lose, and no amount of accuracy can enable it to hold the authority to decide. Nothing in this part depends on anything in the last one, although we hope the former sets an intellectual foundation, or at least an intellectual frame, for what’s to come.
VII. What a license is
Everything to this point has largely involved arguments about capability — with some philosophical ramblings in between — and arguments about capability tend to age poorly. They are hostage to the next release, the next benchmark, the next paper.
In short, were our case for governed clinical execution resting on the bare claim that models cannot reason in language, it would be somewhat arguable today and might, though we doubt it, be in ruins in the future.
Notwithstanding the academic throat-clearing above — which we believe was worthwhile nonetheless — the question that has organized us so far, whether machines can reason in language, is for clinical purposes the wrong question, or at least not extended appropriately.
Because a clinical decision is, before anything else, an act. It is performed, for a particular patient, at a particular hour, by a professional who is answerable for it — answerable in the concrete form with attendant consequences such as litigation, dismissal, and professional discipline, all of which run through the license she holds to be able to act at all. The license was granted for her expertise; the expertise, on its own, entitles her to nothing.
Thus, answerability is not incidental to the license; it is what the license is. A license is a standing liability, and it can be revoked. This a model does not have and cannot acquire, because accountability presupposes something to lose, and even in a world in which litigation can be against the model or organization that created the model, the kind of accountability that for centuries patients have sought from malpractice reaches something deeper, and more human, than just monetary damages. As such, passing a medical licensing examination does not make a model a physician; it simply makes it a model that passes examinations.
Put another way, capability is an empirical property of a system, while authority is an institutional or statutory fact about who is entitled to decide and who bears the cost of deciding wrongly, and no quantity of the first automatically converts into the second.
The question — how capable must the model become before it may decide? — has no answer because it has no referent. The model does not decide — at least until the model is so capable that it is granted institutional or statutory authority ipso facto. Thus, the question worth asking runs orthogonally: what does the clinician need in order to recommend well?
And the question has an answer — rooted in statutory language, now, but also that would strike most as self-evident: a clinician cannot exercise judgment over a recommendation whose basis she cannot inspect and understand the derivation of (ideally, as quickly as possible).
To inspect a recommendation is to be able to see which facts were used, which criteria fired, which version of which standard was applied — and therefore to be able to disagree specifically: that the family history has been misread, that the standard says one thing but this patient represents a more nuanced situation, etc.
A probabilistic draw from a language model affords her none of this. A fluent narrative can be accepted or rejected wholesale, on impression, but it cannot be inspected or interrogated, because there is nothing inside it to inspect or interrogate except, perhaps, a seemingly ceaseless chain of what’s called thought (or reasoning) in a linguistic sleight of hand by the labs.
True inspection, in clinical matters, requires an object while deference requires only a voice. As such, authority assumed without a basis for inspectability is not true authority — it is a verdict that cannot be checked against a standard and cannot be interrogated by those entitled to do so.
In the United States this has been hardened into statute. Under §520(o)(1)(E) of the Federal Food, Drug and Cosmetic Act, as the FDA interprets it in guidance finalized in 2022 and revised in January 2026, clinical software escapes regulation as a device only if it meets four criteria, the fourth of which requires that the software let the health-care professional independently review the basis for its recommendations, so that she does not rely primarily upon them in making a diagnosis or treatment decision for an individual patient.29
To satisfy this, the agency expects the software or it’s labeling to state, in plain language, the intended use, the intended population, the required inputs, and the development and validation of the underlying logic. The reasoning behind the rule is the reasoning above: software whose basis cannot be independently reviewed is not supporting a licensed decision-maker; it is attempting to be one.
The agency names automation bias directly, as we elaborate on below, noting that it breeds both errors of commission, where wrong advice is followed, and errors of omission, where the clinician fails to act because nothing prompted her to. A language model may in principle satisfy the criterion, if it genuinely enables independent review of its basis; the line is drawn not against language models generally but against unexaminable ones specifically. In practice, for a fluent probabilistic system under time pressure, those two categories collapse into each other since it’s language all the way down (including vis-à-vis models that lean on their so-called ‘reasoning’).
The FDA’s analysis turns on two properties of the software together: the level of automation, and the time-critical nature of the decision. Software that hands the clinician a single selected output, in a setting that demands an immediate response, will generally fail, on the ground that she has neither options to weigh nor the time in which to weigh them.30
This has been read as an impasse because it describes front-line practice (e.g., the typical twelve-minute consultation) and taken at face value it appears to condemn clinical software informed by state-of-the-art models to the pixelated dustbin.
But this seeming impasse is rooted in an erroneous assumption: that the review in toto must be performed by the clinician, in the room, at the moment of use. Nothing in the statute requires this, and nothing in the practice of any other profession would suggest it. The engineer does not test the steel in the moment she signs the drawing. The steel was tested long before it reached her to a standard the engineer understands and appreciates as sufficient.
The FDA’s requirement, properly understood, is that the basis of an output be reviewable. Software satisfies the criteria when the logic it executes in a given clinical setting has already been examined and approved by the body that governs the standards in the jurisdiction where it will be used; when what arrives at the point of care is therefore defensible in itself rather than defensible only in retrospect; and when the audit record exists to be consulted should anyone wish to, including the clinician in the moment, but without that inspection being a precondition of safe use.
The clinician who wishes to interrogate a compiled artifact — as created by ourselves, for example — that informs a recommendation shown to her can do so in seconds, because it is at base a rule encoded in logic presented in a way that’s quickly intuitable to a clinician.
And the clinician who does not take the time to interrogate the logic herself is nonetheless working with logic that her governing body has already interrogated on her behalf — and, being deterministic, that logic will not produce for her the scattershot outcomes that a sampled or probabilistic model used alone would.
Canada, where we have implemented our deterministic logic in clinical settings, has largely drawn the same line as the FDA in different words. Health Canada excludes decision support from device regulation on four criteria, three of which track the American ones closely — down to naming clinical guidelines among the medical information such software may display.31
The fourth asks not whether the software enables the clinician to review the basis for its recommendations but whether it is intended to replace her judgment in making the decision. On its face that is the gentler question, since it asks nothing about review; in practice, for a system that cannot be examined, the two questions collapse into each other. A recommendation whose basis cannot be inspected leaves judgment nothing to work on — the clinician can accept it or override it on instinct, and a tool that leaves her only those two moves has taken the decision over in effect. It thus falls foul of the Canadian criterion on its function, and the review requirement the wording omits comes back in through the word replace.
Nor is the federal criterion the only Canadian rule such a tool must survive; the sharper ones are provincial. The colleges that license the country’s clinicians have begun writing their expectations for these tools down, and where they have, the advice is unhedged: the clinician is accountable for her use of them, in clinical decision-making as much as in documentation. She is to review everything they generate for accuracy and completeness and, no law specific to the technology yet existing, the expectations of the profession are unchanged.32
Those expectations can still be met through the use of a compiled artifact because the logic that will run, being deterministic, is set in stone and can be read in seconds: which facts, which criterion, which version, by the person whose responsibility it is to know it. By contrast, of course, this same approach cannot be undertaken against a typical LLM output, whose basis will not survive to the next run and whose logic perhaps will differ even if the ultimate outcome remains the same (never mind what a chain of thought looks like and its incoherent changes prompt-to-prompt even if the output remains the same).
What the colleges have written, without quite setting out to, is a specification, and what it compels is the use of deterministic logic in clinical decision making. This is the only construction under which AI-enabled clinical software that informs decisions is feasible at all, and it is achievable — as we can attest, since we have put it into practice ourselves.
VIII. The deference problem
There remains the possibility — and it is the industry’s talking point du jour — that capability will dissolve the problems we’ve touched on: that a model accurate enough will render the whole question of inspection moot, on the basis that one does not inspect or audit what one has no reason to doubt.
Beyond the numerous statutory and regulatory difficulties — to use a euphemism — with this line of thought, the literature on automation bias says otherwise. The dominant failure mode of clinical decision support has never been that clinicians distrust the system; it is that they defer to it.33
And that deference is not irrational. The relationship between measured accuracy and human scrutiny is monotonic, and it runs the wrong way: the better a system performs, the more sensible it becomes to accept what it says, and the less often anyone examines whether accepting it was sensible on this occasion.
This point is worth putting with some care because it’s one of the rationales behind the statutory and regulatory developments we’ve discussed — and the reason these developments won’t be rolled back no matter how ‘advanced’ probabilistic AI models become (if anything, they will tighten).
Let p be the probability that a given recommendation is correct, and let q be the probability that the clinician, reviewing independently, catches an error when one has occurred. Deference is the observation that q is not a constant but a declining function of p: the more reliable the model, the thinner the scrutiny. The direction of that decline is what the human-factors literature reports; its shape we stipulate, since no one appears to have fitted the curve. The rate at which an error reaches the patient is then
The naive expectation is that r falls as p rises — in other words, that better models mean fewer harmed patients, monotonically, for ever. Differentiating shows the expectation to be unsafe:
Whenever deference erodes quickly enough — whenever the scrutiny lost to rising confidence outweighs the errors removed by rising accuracy — residual harm increases with capability. The sign is the whole of the surprise: harm goes up as the model gets better, while everyone involved behaves reasonably — the clinician’s deference rational, the model’s accuracy real — and the arithmetic is what it is. This isn’t a pathology of bad systems but rather an essential attribute of good systems used without inspection.34
A compiled artifact (a clinical standard, compiled into deterministic logic) breaks this mechanism at its root. Because the basis can be interrogated at low and roughly constant cost, scrutiny no longer decays with confidence; q holds at some floor q₀, and residual harm resumes falling monotonically with accuracy, as everyone had assumed it was doing all along. The artifact instead restores the relationship between accuracy and safety that deference quietly destroys.
MODEL, NOT DATA · The direction of the decay is what the human-factors literature reports. Its shape is stipulated, not fitted — no one appears to have measured the curve. The claim is qualitative.
inspect
| parameter | value | status |
|---|---|---|
| q max — peak catch rate | 0.70 | stipulated |
| p₀ — decay midpoint | 0.85 | stipulated |
| k — decay rate | 2–60 | the slider |
r(p) = 1000·(1 − p)(1 − q(p)), with q(p) = q max / (1 + ek(p − p₀)). The direction of decay is reported; the shape is not fitted — see n. 31.
In the end, increasing capability of probabilistic models does not relieve the pressure for inspectability; it intensifies it. A more persuasive model makes independent review harder, not easier, unless the recommendation arrives as something the clinician can inspect in full.
However, if the output is driven by a compiled artifact — these facts, these criteria, this rule, this version — she has something to think against. She can name the disagreement and overrule the artifact on patient-specific grounds which is the whole of what professional judgment consists of. Whereas, if the output is an eloquent paragraph, self-certain in its pronouncement, her only available moves are to defer based on the ephemeral chain of thought or to override on instinct. The human-factors literature is unambiguous about which of these a clinician with four minutes left in the consultation must choose.
IX. The missing certificate
The compiled artifact has been offered as a guard against the erosion of judgment — as the condition under which a clinician’s authority remains real rather than ceremonial. This is, indeed, what such an artifact prevents but we’ve elaborated little on what it also provides.
The affirmative case for artifacts in the clinical domain is the more interesting vantage point, and it begins by looking, without a slant, at the position the licensed clinician actually occupies.
In short, a clinician is answerable for the correct application of knowledge no human being could hold at command. Clinical standards are written to be read, not executed: long prose documents, revised on their own schedules, diverging across national, regional and institutional pathways, and rule-heavy in ways that ordinary memory cannot survive.
What comprises a standard is not a mystery, and finding the appropriate standard is not hard. What is hard — what is, in a twelve-minute consultation with a patient in the room, close to impossible — is applying standards correctly, in real time, across dozens of clinical domains, and then doing it again for the next patient, and the next, and the one after that.
The liability is total and the epistemic support is thin. As such, the predictable equilibrium is defensive practice: the quiet retreat from best to safe, which is less a failure of character than what any rational professional does when she bears the entire downside of a judgment for which she has been given little firm ground.
Every safety-critical profession has met and addressed this kind of problem before. For example, the structural engineer signs a design and is answerable for it. She does not assay the steel. The steel arrives with a mill certificate — a grade, a heat number, a schedule of guaranteed properties, certified against a standard she had no hand in writing and does not propose to relitigate at her desk. If the steel is not what the certificate says it is, the failure is attributable, upstream, to the mill. The certificate does not diminish her responsibility for the design by a single degree. Rather, it is the condition under which she is able to discharge it.
Likewise, the clinician prescribing a drug is answerable for the prescription. She does not assay the compound. It arrives under a monograph she did not write, approved by a body on which she does not sit, against a manufacturing standard whose enforcement is somebody else’s statutory problem. Nobody has ever regarded this as an erosion of clinical authority. It is what renders clinical authority practicable at all since, of course, it would not be possible to practice medicine for a single afternoon if every physician had first to verify every molecule she prescribed.
These are liability layers, and their essential property — the one on which everything here turns — is that they are additive rather than substitutive. They do not move risk off the professional; they give her warranted ground on which to bear her actual risk. The signature at the foot of the page still means what it has always meant. It now rests on components that are separately guaranteed, separately traceable, and separately correctable when they fail.
Medicine has built these layers for the molecule and for the device — but it has never built a layer for reasoning. The logic by which a standard of care is applied to a particular patient remains, at present, compiled in the mind of the clinician at the bedside — from memory, under time pressure, without provenance, without version, without any certification at all. It is the last artisanal component in a safety-critical chain, and it happens to be the component through which nearly every clinical decision must pass.
What a compiled, verified, inspectable artifact supplies is that missing certificate. Its claim is narrow and it is checkable: on these facts, this criterion, of this standard, at this version. It does not diagnose, treat, or sign, and it is not answerable to the patient — it cannot be, for the reasons already given.
What it does is arrive with its attestation attached, on the proof-carrying discipline of §V: one accepts the artifact not because its generator is trusted, but because the evidence travels alongside it, will not produce a different answer when the same facts are presented, and may be checked by someone who trusts nothing and nobody.
The clinician who accepts such an artifact-derived recommendation accepts it on grounds. The clinician who departs from it departs on grounds, and her departure is itself a documented act of professional judgment rather than an unexplained divergence. In both directions her position is strengthened, and in neither is any part of her authority given away. This is the sense in which the artifact enhances the license rather than encroaches upon it: not by taking any of the burden, but by giving her something to stand on while she carries it.
To be clear, an artifact of this kind does not indemnify the professional and makes no promise whatever about the outcome for the patient. The artifact promises only about itself: that it is what it says it is, that it did what it says it did, and that its reasoning can be found and inspected by anyone who cares to look and has been accepted by the appropriate governing bodies or institutions. In short, an artifact can warrant only itself; the professional warrants the decision.
The two liabilities at play, properly constructed, run in different channels and neither absorbs the other. The clinician’s liability flows through the license: to her, personally, discharged by the exercise of judgment and enforced by the courts and the licensing college. The artifact's liability flows through certification and warranty: to its maker, discharged by fidelity — the artifact correctly compiled and approved by a governing body — and enforced by contract. Both channels are live. Neither is a substitute for the other, and a system properly built has an interest in keeping both open.
Which is what a black box — such as a modern frontier model — cannot emulate, and this is the deepest objection to it. A black box does not constitute a liability layer; it constitutes a liability sink.
inspect
| channel | basis | enforced by |
|---|---|---|
| Clinician | license — a standing liability | courts and college |
| Warranted artifact | certification — it is what it claims | regulator and contract |
| Black box | no attribution possible | — exposure falls on the license-holder |
Static comparison; arrow weights are illustrative, not quantified. Construction follows §IX: two channels, neither absorbing the other; remove the artifact's channel and the exposure does not disappear — it concentrates, with interest. The artifact node carries the same version and hash as Figs 7–8; the repair footline follows §IX — what can be attributed can be corrected.
It cannot bear attribution, because attribution requires inspection and there is nothing there to reliably inspect. So, when its recommendation is wrong, the failure has nowhere to land but on the professional — who accepted it because she had no means of interrogating it, and who is now answerable for a determination she was never in a position to make. It subtracts confidence and adds exposure, which is the inverse of what a supporting layer does. This is the automation-bias problem seen from its other side. She defers because she has nothing to argue with — the voice is perfectly calibrated, and there is nothing behind it — and having deferred, she stands alone.
It will be objected that state-of-the-art models are not black boxes at all — that they show their working, that a chain of thought may be printed out and read, and that the J-lens can now read even the reasoning they decline to print. To which the answer is: run it all again, on the same facts — is the chain, in the general use of the word, identical?
If it is not, then the chain was never a basis. It was a narration, composed afresh on each occasion, and its relation to the recommendation is not that of a reason to a conclusion but that of a caption to a photograph. One cannot audit such a thing, because the audit does not reproduce the object under audit; one cannot correct it, because there is nothing stable to correct; and one cannot rely upon it, because reliance presupposes that the thing relied upon can be replicated tomorrow. A basis that varies under re-execution on identical inputs is void — not weak evidence, but no evidence — and this is true however lucid the prose, and however deep the purported reasoning underneath it.35
The objection therefore founders on a distinction: legibility is not warrant. To see what a model was thinking is a fine thing, and the interpretability researchers are right to pursue it. But a professional who must answer for a specific, fact-based decision does not need to know all of what the model was thinking. She needs to know what it will do, and that it will do it again, and that the grounds on which it did it can be produced, unchanged, in a courtroom four years from now. A properly compiled artifact meets that standard by construction. A sampling distribution does not meet it at all, and no level of compute thrown at it will make it do so.
This brings us to abstention which has a precise formulation belonging to the theory of warrant. A system that answers selectively is a pair — a predictor and a gate that decides when to answer at all. Its coverage is the fraction of cases it accepts, and its selective risk is the error rate on the cases it does not decline:36
The guarantee that makes the construction worth having is that, for any target risk, one may hold the selective risk below it by spending coverage: the system declines the cases in which confidence is low. Abstention, so understood, is the decision that purchases the guarantee.
For example, when a medical laboratory receives a hemolyzed specimen it does not estimate the potassium. It reports the specimen unsuitable and requests a redraw, and no clinician anywhere has ever regarded this as a failure of the laboratory. It is the behavior that makes the laboratory worth relying upon in every case where it does report.
I do not know; escalate is therefore not an admission of weakness in a clinical system. It is the system keeping the only promise it ever made — the promise about itself — and returning the determination to the party entitled to make it.
A skeptical clinician’s first objection to all of this talk of determinism and artifacts is that her standards are not always trees — that some standards contain elements written in the conditional, full of ‘consider’ and ‘may be offered’ and ‘in selected patients, after discussion of risks and benefits’, and that no honest compilation of such documents could produce a determination at every leaf.
She is right, and the objection must be reflected in the design. Where the standard commits, the artifact commits; where the standard defers to judgment, a faithful compilation compiles the deferral — an explicit leaf that names the factors the standard names and hands the determination to the person the standard hands it to. The conditional mood is not a defect to be engineered away; it is part of what the standard says (sometimes), and an artifact that resolved it on the standard’s behalf would have failed at the only thing it exists to do, which is to say what the standard says.
Part Three
The Same Design
The dispute of §VI, left open, turns out not to need closing. Its two branches arrive at the same architecture — which wants the most capable models that can be built, to be built, but wants them at some remove from the consulting room.
X. Where the intelligence belongs
An argument of this shape invites a misreading which should be closed off before it hardens. Nothing said here is an argument for less artificial intelligence in medicine, or for weaker models, or — heaven forbid — for a return to the hand-built expert systems of the nineteen-eighties, whose failure was total and deserved.
Our position is the opposite. The architecture we propose desires the most capable models that can be built to be built. It simply wants them to have an indirect impact on the consulting room.
The distinction is between build time and run time. At build time the frontier model does what only a frontier model can do: it reads a vast corpus no individual has ever read in full — the guidelines, the updates, the evidence tables, the jurisdictional variants — and proposes the logic latent within them (this criterion, these thresholds, this branch, this exception).
inspect
| stage | actor | object | standing |
|---|---|---|---|
| Build | frontier model | candidate | probabilistic · no authority |
| Adoption | the body | reviews and approves | the only step that confers authority |
| Run | — no model | compiled artifact, versioned | deterministic · same facts, same output |
| Rendering | static native SVG · final state shown |
Construction follows §X. The artifact is the same object on both sides of the gate; the model is not.
This work is the work described in §V. A stochastic, high-capacity, linguistically fluent process searches an enormous space and proposes candidates, as FunSearch proposes functions and AlphaGeometry proposes constructions. It is the ideal proposer for the task, and — this is the part the argument has been building toward — the better it becomes, the better the task is done.
What the model drafts is not an artifact; it is an artifact candidate. It becomes an artifact after it’s been fine-tuned (as we do) and when the authority entitled to approve it in a jurisdiction — the college, the society, the health authority, whichever governing body already owns the standards setting — has examined it, amended it as needed, and approved it. Nothing that has not been approved is ever used. The model never has the last word, upstream any more than downstream.
The division of labor between the compiler and the governing body is exact. What the compiler can prove about a compiled artifact — before any patient is involved, and in the strict sense of prove — is that the logic is whole: every combination of findings leads somewhere; no two rules fire in contradiction; no branch sits unreachable.
However, what no compiler can prove is that the logic is the standard’s — that the tree, however sound, says what the standard is meant to say. The final judgment belongs to the people who hold the standard — which is why the governing body’s review is not a formality performed over the model’s work but the only step in the pipeline that confers any authority at all.
This is not a novel governance arrangement. It is the arrangement medicine already has, and has had for a century, alongside most other sensitive or established sectors. The engineer does not ratify the grade at her desk; the clinician does not ratify the monograph at hers, and nobody has ever asked either to. Each is entitled to interrogate the standard that governs her, and each inherits it on the authority of a governing body that examined it on her behalf and can be held to account for having done so.
What is being automated here, then, is not judgment. It is the labor of codification — the reading, the cross-referencing, the reconciling, the drafting — which is the work that has never been done at scale in medicine, for the obvious reason that doing it by hand was never feasible with interconnections so vast.
It will be asked, fairly, whether this review is not subject to the very automation bias we have just made so much of — whether a governing body approving an artifact candidate does not defer to it as the clinician at the bedside would.
It does not, and the reason is the reason already given. Approving is the deliberate acceptance of a finished artifact candidate, not the real-time acceptance of a probabilistic output. The work is conducted on a compiled artifact with its provenance attached, not on a probabilistic sentence.
And the artifact candidate is given to several reviewers, checked against the sources, and argued over, all before it reaches a clinician to use — which is to say it is subjected to the adversarial pressure that a clinician with a waiting room cannot apply. Then, with the finalized artifact validated and versioned, changes can be made as the standards change and be corrected everywhere at once.
Moving the model's contribution into this setting is what makes it safe since the model itself is never used at run time. The danger was never the model's help; it was the model's help arriving where there was no time to weigh it.
Now, it should be said, a large language model that sufficiently enabled independent review of its basis could satisfy the fourth criterion as laid out by the FDA and Health Canada; the line is drawn against unexaminable software, not against large language models as such.
This, for all practical purposes, we decline based on the argument of §VIII rather than on any squeamishness about the technology: a system whose safety depends on being inspected in the moment, by a clinician who has no moment, is a system whose safety depends on a review that will not happen.
With all that said, let’s suspend disbelief and suppose a model reasons at the point of care but every step is checked as it goes — a proof-carrying chain, verified in the room, the certification work of §V brought to the bedside. Given this, what remains to object to?
To return to a prior refrain: run it again. The checker will happily certify a different chain, since what it certifies is the validity of the steps taken on this occasion, and a family of individually flawless derivations that disagree (or, put differently, are presented in a different light) is the sampling distribution of §IX wearing a different gown.
The certificate attaches to the run, and the run cannot be produced again; what the courtroom is handed four years on is an attestation that something valid seemed to occur. The clinician, meanwhile, in the moment is handed something worse — the most deference-inducing output yet devised, an eloquent recommendation with a metaphorical (or literal) checkmark on it.
Thus, what per se is required in clinical settings, in whatever conception or format one wishes to dream up, is an artifact: versioned, deterministic, and identical on Tuesday to what it was on Monday. The same facts produce the same output, and will produce it again in four years, and can be produced in a courtroom. Nothing is sampled. Nothing is inferred afresh. The clinician sees which rule fired and on what facts, and departs from it or does not, on grounds. Asked why, she has an answer that is not the machine said so.
the facts
Age
Findings
the consultNO MODEL · NOTHING SAMPLED
—
inspect
| clause | criterion | points |
|---|---|---|
| mc-age-3-14 | age 3–14 | +1 |
| mc-age-45 | age ≥ 45 | −1 |
| mc-exudate | tonsillar exudate | +1 |
| mc-nodes | tender anterior cervical nodes | +1 |
| mc-fever | fever ≥ 38 °C | +1 |
| mc-no-cough | absence of cough | +1 |
| score | directive |
|---|---|
| ≤ 0 | No testing; no antibiotics — symptomatic care. |
| 1 | No routine testing; reassess if the course is atypical. |
| 2–3 | Rapid antigen test; treat only on a positive result. |
| ≥ 4 | Rapid antigen test; consider empiric therapy while awaiting the result. |
These two tables are the artifact. The function in this page walks them and does nothing else. Clause ids follow the candidate PlanDefinition; the scoring is the McIsaac modification of the Centor criteria.
ILLUSTRATIVE · simplified from a PCare candidate artifact (PlanDefinition plan-strep-triage, v0.1.0; CQL library cond_test). Not medical advice, and not the shipping product — the deployed system compiles from adopted guidelines with full provenance. The rule shown is complete under inspect.
Versioning, while we are here, is usually described as an audit convenience but it is undersold. Standards move; the task forces that write them exist to move them. When a task force lowers a starting age or retires a threshold, a versioned artifact does not update in place — it records: what changed, when, on whose approval, and which patients were assessed under which version on which day. The diff between version four and version five is itself a clinical document and, by extension, enables an automatic reassessment of prior patients who went through the former standard based on the latest standard.
Ask a frontier model deployed at run time what it would have recommended last March and there is nothing to ask — the weights have moved, silently, and the March model no longer exists to be interrogated unless it’s somehow rolled back.
This is what §VII promised vis-à-vis our proposed architecture. The basis is reviewed before the consultation rather than during it — once, by the governing body entitled to set the standard, on behalf of every clinician who will subsequently rely upon it — and so the twelve minutes are no longer an obstacle to review, because the review has already happened. Defensibility is a property the artifact carries into the room, not a burden it imposes on whoever finds it there.
For a system that reasons at run time, an improvement in the underlying model is a mixed blessing: the answers get better, the deference gets deeper, the automation bias gets worse, and the warrant does not improve by a hair — because there was never any warrant there to improve. For a system that compiles at build time, an improvement in the underlying model is pure gain. More of the corpus is compiled, more quickly, more cheaply, more thoroughly — and not one increment of new risk arrives at the bedside, because the bedside never meets a model. Frontier progress is a tailwind with no answering headwind.
There is, in consequence, no version of the dispute in §VI in which better models make this architecture, as we’ve described it, worse. If reinforcement learning merely sharpens sampling, the compiled artifact is what supplies the guarantee that sampling cannot. If reinforcement learning genuinely expands machine reasoning, then the proposer improves, the corpus is compiled faster and better, and the guarantee is unaffected — because it never depended on the proposer in the first place. One may be as bullish about artificial intelligence as one likes, and arrive at the same design.
Which brings us to trust, and to the reckoning that is plainly coming. Public and professional confidence in clinical artificial intelligence is not robust, and every hallucinated citation, every confidently wrong recommendation, every case in which a clinician is asked to explain why she followed a model she could not interrogate, draws further on a diminishing account. The technology industry’s standing answer is to request more trust: trust the model, trust the evaluation, trust the guardrail, trust us.
This architecture we propose requests none of this, because it puts no model at the point of care. It is artificial intelligence in its construction and not artificial intelligence in its operation — a system that could not have existed before the frontier models but crucially does not depend upon the frontier models when it matters most. The clinician is not asked to trust; she is asked to check, and she is able to. Do not trust; verify is the only posture that survives a backlash, and it is also the only posture that has ever been available to a profession that must answer, personally and in public, for what it does.
With this, the double asymmetry of §I is now discharged. The better the models become, the more urgently medicine requires that its own standards be executed deterministically — because capability, as §VIII showed, buys deference before it buys safety. And the better the models become, the better those compiled artifacts become — because the model that drafts them is the model that improved. The two halves are not in tension. They are the same fact, observed from the two ends of the pipeline.
The apparent paradox is not one. The approach is possible only because the models have become exceptional, and it is safe only because it declines to let them into the room. Both halves are necessary. It requires the frontier to be excellent, and it requires the frontier to stay out. And by doing so, clinicians are empowered to do far more work, in far less time, with far greater confidence that they’re making the right decisions.
XI. Either way
In the end, the capability dispute of §VI is unresolved — and may remain so for years to come — but it does not matter.
Suppose Yue and colleagues are right, and reinforcement learning merely sharpens a model’s aim within a distribution it already possessed. Then the output of even the most capable reasoning model remains a draw from a probability distribution — more skilfully aimed, but a draw nonetheless, and never a certificate.
It is not reproducible, since the same input under sampling may yield a different chain and a different answer. It is not auditable, since the winning trace is a (debatably) persuasive narrative rather than a checkable proof object. And it does not reliably abstain, since a system that reasons brilliantly across five problems does not always thereby know, on the sixth, that it should decline. A warranted, inspectable, versioned substrate is required for reliability.
Suppose instead that Liu and colleagues are right, and prolonged reinforcement learning (or, we would add, any other RSI approach) genuinely expands the reasoning boundary, and the models of the coming decade are more capable, more accurate, and more fluent than anything now in service. Then the clinical problem at the point of care is not solved but aggravated.
The more persuasive the system, the greater the pressure toward deference; the more accurate the system, the more rational that deference becomes; and a clinician confronted with an eloquent and unexaminable recommendation, in a consultation with four minutes left in it, has no purchase from which to disagree. Her licensed judgment survives on paper and evaporates in practice. A warranted, inspectable, versioned substrate is required — more urgently in this scenario — for authority to remain.
The branches converge, and they converge because our clinical argument never depended on the models being weak. It cannot, therefore, be defeated by the models becoming strong. It rests on four propositions that no benchmark can disturb: that the determination belongs to a licensed professional who is answerable for it; that nobody can exercise judgment over a basis she cannot inspect; that a basis which will not survive re-execution on identical facts is not a basis at all; and that an output which cannot be inspected cannot be blamed, and therefore cannot be relied upon, and therefore cannot be used.
The architecture that satisfies all four is available. A high-capacity, stochastic, linguistically fluent process may propose, may search, may draft, and may explain; a reproducible, inspectable, formally checkable artifact carries the logic that reaches the decision, and carries its warrant with it; and the decision itself remains, in every case, the clinician’s to make and the clinician’s to answer for — but answered for, at last, on ground that is carefully documented and approved by the clinician’s governing body. The regulatory framework, read closely, already requires this, whether anyone fully recognizes it or not right now.37
Coda
What the neuroscience contributes, in the end, is not a prohibition but a precedent. It was never going to demonstrate that a machine cannot reason in language; it is an existence proof, not an impossibility proof. What it demonstrates is that general, abstract, formal reasoning can run on a substrate that is not natural language — that the linguistic surface and the inferential engine are separable in principle because in the one uncontested instance of general intelligence available for inspection they are in fact separated.
Evolution did not build a fluent system and hope that valid inference would appear somewhere inside it; it built the inference elsewhere and put language at the interface, where language belongs, doing what language is for — as the brain, which has been running this architecture for a rather long time, has been demonstrating all along.
In short, language is the interface, not the engine, and we expect this to hold true in a fundamental sense for frontier models even as they continue to advance. However, whether our view on this ongoing argument is correct is irrelevant in the clinical domain we operate in because a clinical decision now, and for the foreseeable future, flows through a license — and that license now, and also for the foreseeable future, must attach to a human and to their sometimes all-too-human judgment.
Given this, AI's contribution to a clinical recommendation must be indirect, or it must be non-existent. The license has always implied as much, and now both American statute and Canadian regulations say so in plain language. Indirect, in this case, means that the frontier model of choice does its work at build time, drafting a standard into an inspectable, versioned artifact candidate; that this candidate is fine-tuned by expert physicians; that the governing body which holds the standard reviews and approves it; and that the clinician, with a patient before her, then relies upon the approved artifact, with the full weight and scope of the applicable standard and the governing body’s stamp of approval behind it — a liability layer added under her signature, never a substitute for it.
This is what the license demands, what the nouveau statutes and regulations demand, and what clinicians and those they serve will always demand.
Provenance
- Version
- 33 · 19 July 2026
- Integrity
- sha-256 · …
- Length
- 14,262 words · 37 notes · about 64 minutes
- Sources
- Verified against original publications; locators to page, table and figure
- Contested
- §VI — conceded without qualification, and left unresolved
- Interest
- The authors build clinical software of the kind argued for here
Cite
Compiled Health, ‘Reasoning Is Not Language’, v33, 19 July 2026. sha-256
This essay makes a claim about artifacts that arrive carrying their own warrant. It seemed only fair that it should do so itself.
Notes
- 1Aristotle, De Interpretatione, in Categories and De Interpretatione, trans. J. L. Ackrill (Oxford: Clarendon Press, 1963); the relevant passage is 16a3–8. ↩
- 2B. L. Whorf, Language, Thought, and Reality (Cambridge, MA: MIT Press, 1956); P. Carruthers, ‘The cognitive functions of language,’ Behavioral and Brain Sciences 25, no. 6 (2002): 657–674; G. Lupyan, ‘The centrality of language in human cognition,’ Language Learning 66, no. 3 (2016): 516–553. ↩
- 3J. A. Fodor, The Language of Thought (New York: Thomas Y. Crowell, 1975). ↩
- 4E. Fedorenko, A. A. Ivanova and T. I. Regev, ‘The language network as a natural kind within the broader landscape of the human brain,’ Nature Reviews Neuroscience 25 (2024): 289–312. ↩
- 5On arithmetic and algebra: M. Amalric and S. Dehaene, ‘A distinct cortical network for mathematical knowledge in the human brain,’ NeuroImage 189 (2019): 19–31; M. M. Monti, L. M. Parsons and D. N. Osherson, ‘Thought beyond language: neural dissociation of algebra and natural language,’ Psychological Science 23, no. 8 (2012): 914–922. On computer code: A. A. Ivanova et al., ‘Comprehension of computer code relies primarily on domain-general executive brain regions,’ eLife 9 (2020): e58906; Y.-F. Liu et al., eLife 9 (2020): e59340. ↩
- 6J. Duncan, ‘The multiple-demand (MD) system of the primate brain: mental programs for intelligent behavior,’ Trends in Cognitive Sciences 14, no. 4 (2010): 172–179. ↩
- 7H. Kean, A. Fung, P. Jaggers, J. Chen, J. S. Rule, Y. Benn, J. B. Tenenbaum, S. T. Piantadosi, R. A. Varley and E. Fedorenko, ‘Evidence from formal logical reasoning reveals that the language of thought is not natural language,’ Proceedings of the National Academy of Sciences 123, no. 28 (2026): e2520095123. The version of record is paywalled; the open preprint (bioRxiv 10.1101/2025.07.26.666979, v3, 6 May 2026, CC-BY) carries the same results, and the locators here are to it. Neither the deductive contrast (Modus Tollens over Modus Ponens) nor the matrix contrast reaches significance in the language network — ps > 0.1, at Figure 2A–B and Table 1 — despite behavioral difficulty effects in both, at SI Figure 4. The inductive contrast is reliable but small: the authors put the language contrast at over four times stronger, and the induction condition itself at or below the nonword-reading baseline, at Table 1 and SI Figure 1B. ↩
- 8The patient data are from the same study, at n. 7, and the locators are again to the preprint cited there. Etiologies, lesion extent, the agrammatism classification and the at-or-near-chance comprehension of reversible sentences are at SI Table 3; the reversible sentence is the paper’s own example. On induction, S.A. solved 19 of 25 rules and G.S. 39 of 40 — rules, each of which comprised several input–output problems — and neither differs significantly from the control sample (Crawford–Howell, ps > 0.499). On WASI-II matrix reasoning, raw scores of 25 of 30 and 26 of 30 place them 2.3 and 1.8 standard deviations above the age-matched mean, which is where the T-scores of 73 and 68 come from, against a normative mean of 50 and a standard deviation of 10: Results, and Figure 2D. That the sample is two patients is the first limitation the authors record, and they give the reason — aphasia this profound is rare. ↩
- 9R. Varley and M. Siegal, ‘Evidence for cognition without grammar from causal reasoning and “theory of mind” in an agrammatic aphasic patient,’ Current Biology 10, no. 12 (2000): 723–726; R. A. Varley et al., ‘Agrammatic but numerate,’ PNAS 102, no. 9 (2005): 3519–3524. These are the earlier patient studies, and they appear as references 22 and 23 in the study at n. 7; the scores given above are from that study, not from these. ↩
- 10R. A. Varley, ‘Reason without much language,’ Language Sciences 46 (2014): 232–244, §7, ‘Limits to the evidence from aphasia,’ where the objections are set out, together with Chomsky’s appeal to competence and H. A. Whitaker’s reply to it. ↩
- 11H. Kean et al., ‘A human brain network specialized for abstract formal reasoning’ (preprint, bioRxiv 2025.10.21.683445). ↩
- 12E. Fedorenko, S. T. Piantadosi and E. Gibson, ‘Language is primarily a tool for communication rather than thought,’ Nature 630 (2024): 575–586. ↩
- 13S. T. Piantadosi, J. B. Tenenbaum and N. D. Goodman, ‘Bootstrapping in a language of thought,’ Cognition 123 (2012): 199–217, and ‘The logical primitives of thought,’ Psychological Review 123, no. 4 (2016): 392–424. ↩
- 14J. S. Rule et al., ‘Symbolic metaprogram search improves learning efficiency and explains rule learning in humans,’ Nature Communications 15 (2024): 6847. ↩
- 15K. Mahowald, A. A. Ivanova, I. A. Blank, N. Kanwisher, J. B. Tenenbaum and E. Fedorenko, ‘Dissociating language and thought in large language models,’ Trends in Cognitive Sciences 28, no. 6 (2024): 517–540. ↩
- 16I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio and M. Farajtabar, ‘GSM-Symbolic: understanding the limitations of mathematical reasoning in large language models,’ ICLR 2025 (arXiv:2410.05229). The NoOp condition is §4.4, where the reported drops run as high as 65 percent across the models tested. ↩
- 17N. Dziri et al., ‘Faith and fate: limits of transformers on compositionality,’ NeurIPS 36 (2023); R. T. McCoy et al., ‘Embers of autoregression’ (arXiv:2309.13638, 2023); M. Nezhurina et al., ‘Alice in Wonderland: simple tasks showing complete reasoning breakdown in state-of-the-art large language models’ (arXiv:2406.02061, 2024). ↩
- 18P. Shojaee et al., ‘The illusion of thinking’ (arXiv:2506.06941, 2025); for the rebuttals, on output-token limits and on unsolvable puzzle instances scored as failures, A. Lawsen, ‘Comment on “The Illusion of Thinking”’ (2025). The re-analysis of GSM-Symbolic referred to in the text is a separate paper: D. Długosz, A. Oliveira and N. Díaz-Rodríguez, ‘The importance of being statistically earnest: a critical re-evaluation of GSM-Symbolic’ (arXiv:2605.28700, 2026). Re-running twenty open-weight models under generalized linear mixed models with per-question random effects, they find the variant effect significant for only ten of the twenty; and they identify a systematic shift toward larger integers in GSM-Symbolic relative to GSM-Base (K–S = 0.12, p < 0.001) which accounts for the significance in roughly half of the remainder. ↩
- 19L. Pan et al., ‘Logic-LM,’ Findings of EMNLP 2023; X. Ye et al., ‘SatLM,’ NeurIPS 36 (2023); T. Olausson et al., ‘LINC,’ EMNLP 2023; G. Poesia et al., ‘Certified deductive reasoning with language models’ (arXiv:2306.04031, 2023). ↩
- 20K. Ellis et al., ‘DreamCoder,’ PLDI 2021; B. Romera-Paredes et al., ‘Mathematical discoveries from program search with large language models,’ Nature 625 (2024): 468–475; D. J. Mankowitz et al., ‘Faster sorting algorithms discovered using deep reinforcement learning,’ Nature 618 (2023): 257–263; A. Verma et al., ‘Programmatically interpretable reinforcement learning,’ ICML 2018; O. Bastani, Y. Pu and A. Solar-Lezama, ‘Verifiable reinforcement learning via policy extraction,’ NeurIPS 31 (2018). ↩
- 21G. C. Necula, ‘Proof-carrying code,’ POPL 1997, 106–119; C. Rudin, ‘Stop explaining black box machine learning models for high-stakes decisions and use interpretable models instead,’ Nature Machine Intelligence 1 (2019): 206–215. ↩
- 22T. H. Trinh, Y. Wu, Q. V. Le, H. He and T. Luong, ‘Solving olympiad geometry without human demonstrations,’ Nature 625 (2024): 476–482; Google DeepMind, ‘AI achieves silver-medal standard solving International Mathematical Olympiad problems’ (2024), with the formal results reported in Nature (2025). The 25 of 30 is the Nature paper’s headline result. The silver-medal system of 2024 was AlphaProof paired with AlphaGeometry 2, an upgraded version rather than the system described in that paper; the reported score was four of six problems, for 28 points. ↩
- 23S. Hao et al., ‘Training large language models to reason in a continuous latent space’ (arXiv:2412.06769, 2024). ↩
- 24W. Gurnee, N. Sofroniew et al., ‘Verbalizable Representations Form a Global Workspace in Language Models’ (Anthropic, Transformer Circuits, 6 July 2026). Stanislas Dehaene and Lionel Naccache, architects of global neuronal workspace theory, contributed invited commentary. The technique relies on a single-token approximation, and the mechanism by which a representation enters the workspace remains unresolved. ↩
- 25DeepSeek-AI (D. Guo et al.), ‘DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning,’ Nature (2025); arXiv:2501.12948. ↩
- 26Google DeepMind, ‘Advanced version of Gemini with Deep Think achieves gold-medal standard at the International Mathematical Olympiad’ (2025). The result was not self-graded. DeepMind’s model was among the first cohort whose solutions were marked and certified by the competition’s own coordinators, to the criteria used for student scripts, and the solutions were published; five of six problems, 35 points of 42. No peer-reviewed account has appeared, and the company’s announcement remains the record — we concede the result on the strength of the certification rather than the announcement. ↩
- 27Y. Yue et al., ‘Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?’ NeurIPS 2025 (arXiv:2504.13837). The pass@k estimator, over n samples of which c are correct, is pass@k = E[1 − C(n − c, k) / C(n, k)] ↩
- 28M. Liu et al., ‘ProRL: prolonged reinforcement learning expands reasoning boundaries in large language models,’ NeurIPS 2025 (arXiv:2505.24864). ↩
- 29Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff (final guidance, 28 September 2022; revised final guidance issued 6 January 2026, re-issued 29 January 2026), interpreting §520(o)(1)(E) of the Federal Food, Drug, and Cosmetic Act. The four non-device criteria, the explicit treatment of automation bias, and the position on time-critical decision-making all appear in the 2026 revision. ↩
- 30The two-factor framing — level of automation and time-critical nature of the decision — and the example of software ‘requiring an immediate response (e.g., within 24 hours)’ are drawn from the 2026 guidance. The position is contested: commentators note that clinicians make time-sensitive decisions constantly, and that the agency supports the concern chiefly by citing a single 2004 paper on automation bias in aviation (M. L. Cummings, AIAA, 2004). The point of citing it here is not that it is beyond dispute, but that even a contested regulatory line runs exactly where the argument of this essay runs. ↩
- 31Health Canada, Software as a Medical Device (SaMD): Definition and Classification (guidance document, in effect 18 December 2019), §2.2, which names ‘clinical guidelines’ among the medical information such software may display, analyze or print. The guidance states that the four criteria ‘should not be interpreted as a rigid set of exclusion factors’ but as ‘a foundation for an analysis to be carried out’ — the reading applied here. One honest asymmetry runs the other way: the Canadian exclusion extends to software supporting decisions by patients and caregivers, a category the American guidance of 2022 declined to carry over from its 2019 draft. ↩
- 32The duties quoted are the Ontario college’s — College of Physicians and Surgeons of Ontario, Advice to the Profession: Using Artificial Intelligence in Clinical Practice (updated August 2025) — the fullest such advice yet published; the Canadian Medical Protective Association, which defends the country’s physicians in these matters, reports that many colleges have issued preliminary guidance urging consideration of accountability, transparency and accuracy, while describing the regulatory framework itself as ‘a work in progress’ and the guidance available to physicians for evaluating such tools as limited: CMPA, The Medico-Legal Lens on AI Use by Canadian Physicians (position paper). ↩
- 33K. Goddard, A. Roudsari and J. C. Wyatt, ‘Automation bias: a systematic review of frequency, effect mediators, and mitigators,’ Journal of the American Medical Informatics Association 19, no. 1 (2012): 121–127; B. Vasey et al., ‘Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI,’ Nature Medicine 28 (2022): 924–933. ↩
- 34The model is illustrative, not a forecast: p and q are treated as scalars and the deference response q(p) is posited rather than measured. Its only claim is qualitative — that a sufficiently steep q(p) makes residual harm non-monotonic in accuracy, which is the automation-bias literature restated. ↩
- 35The point is not merely theoretical. Run-to-run variation in language-model inference arises from batching and floating-point non-associativity as well as from sampling, and is not eliminated by setting the temperature to zero; see Thinking Machines Lab, ‘Defeating nondeterminism in LLM inference’ (2025). ↩
- 36C. K. Chow, ‘On optimum recognition error and reject tradeoff,’ IEEE Transactions on Information Theory 16, no. 1 (1970): 41–46; R. El-Yaniv and Y. Wiener, ‘On the foundations of noise-free selective classification,’ JMLR 11 (2010): 1605–1641; Y. Geifman and R. El-Yaniv, ‘Selective classification for deep neural networks,’ NeurIPS 30 (2017). ↩
- 37The lifecycle and governance frameworks point the same way: K. Lekadir et al., ‘FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare,’ BMJ 388 (2025): e081554; and FDA/Health Canada/MHRA, Good Machine Learning Practice for Medical Device Development: Guiding Principles. ↩
