Construct Capture
On benchmarks, specifications and the power to decide what counts
8 SEPTEMBER 2026 · ABOUT 26 MINUTES · BY CHRISTIAN CHARTIER
Contents
OpenEvidence announced a perfect score on the United States Medical Licensing Examination; OpenAI reported that GPT-5.4 inside ChatGPT for Clinicians led HealthBench Professional; Doximity, citing NOHARM, announced that its assistant had outranked OpenEvidence and several frontier models. All three claims can be accurate, an accommodating state of affairs made possible by differences in the tasks, the systems and the ways of marking an answer. The difficulty begins when the results are gathered under clinical AI, as though these separate achievements had settled which system a physician should rely on and what she should be entitled to ask it to do. The reader must supply the relationship among them, presumably on the understanding that somebody has already worked it out.1
The numbers spare the reader considerable work, much of which, one hopes, has actually been done. They imply that the argument happened elsewhere, among people qualified to have it, and that only the result has been forwarded to us. Restoring the particulars produces a less convenient proposition: this system met these cases, with these tools and instructions, and was interpreted by this grader under these criteria. Change the arrangement and another system may win. That is a useful finding, provided its qualifications survive the journey into the hospital, including the meeting in which somebody wants to know what to buy and has, understandably, made limited provision for a seminar on measurement.
At that meeting, three questions need answers: what was assessed, why the expected answers were considered correct, and how the result bears on the proposed use. The middle question is especially easy to overlook, since even a carefully bounded assessment of whether a model follows medicine requires someone to decide what the relevant medicine demands. Once institutions rely on the answer key, those who control it have begun to influence acceptable clinical conduct. We would like to know who made those decisions, what evidence supports them and what happens when a clinician has grounds to say that one of them is wrong.
I. What a score establishes
A benchmark combines tasks, a marking scheme and a procedure for turning the marks into a result; the choice of each helps determine what can be claimed for the whole. HealthBench Professional uses physician-authored consultation, documentation and research tasks, with criteria specifying desirable features and penalising undesirable ones, and a model judges whether a response satisfies them. Its original leading score, 59.0, belonged to GPT-5.4 inside ChatGPT for Clinicians. The prompts, tools and other software around the model were part of the tested product. Giving the underlying model sole credit would therefore change the object described, even while leaving every reported number intact.2
The number also needs its proper unit. A pulse of 59 is measured in beats per minute; HealthBench Professional’s 59.0 consists of length-adjusted, weighted criterion points, averaged under its procedure. An answer has order, proportion, tone and a reader; the rubric reconstructs these from weighted pieces. The resulting index is a legitimate way to organise an assessment, although its interpretation still depends on what earned or lost the points. An omitted qualification and an omitted emergency escalation can both reduce a score. The patient has considerably more reason to care which one was omitted than the percentage, by itself, permits her to know.
Anyone who writes should recognise the problem. A copy editor can identify solecisms, a fact-checker can verify references, and readers can prefer one essay to another; combine their judgments and the arithmetic may be impeccable while the writing it rewards lies quite dead on the page. What each quality contributes depends on what the writer is trying to do, to whom, and with what success—an arrangement in which the author of the present essay has, admittedly, an interest. Medicine likewise contains readily assessable parts, including dose calculations and the recognition of contraindications, alongside a practice involving diagnosis, treatment, continuity, explanation and restraint. A scalar called clinical intelligence has already decided how these achievements contribute to the whole, including when excellence in one can compensate for failure in another.
Psychometrics has been attending to this difficulty for considerably longer than clinical AI has existed. The Standards for Educational and Psychological Testing locates validity in the evidence supporting a particular interpretation of scores for a proposed use; Michael Kane makes the assumptions joining performance to interpretation and use explicit enough to be challenged. Reproducing a ranking can establish consistency while leaving its proposed meaning unsupported. Reliability is a virtue of measurement; validity belongs to the claim made with it. The capacity the assessment claims to measure is its construct: omitting follow-up from a test used to claim competence in patient management is construct underrepresentation; rewarding verbosity in a claim about diagnostic judgment may introduce construct-irrelevant variance. The terminology gives familiar objections a precise address, so that a disagreement about what the test establishes can be pursued through its construction.3
The choice of cases contains a view of medicine too. A harm-weighted benchmark overrepresents dangerous transitions; a prevalence-weighted one fills itself with ordinary care; a frontier benchmark selects yesterday’s failures. None is simply medicine. Each can provide useful evidence, just as an artificial crash test can reveal something important about a collision. NOHARM’s contribution is to examine potential clinical harm directly; in its July version, more than four-fifths of severe errors were omissions. Such failures must remain visible however impressive the other answers, because averaging across them settles how much ordinary success will compensate for one dangerous omission. The arithmetic is performing a clinical judgment, with the additional convenience that it can look like a matter for the statistician.4
There is also the question of whose performance we have measured. In August, a JAMA Perspective led by Ezekiel Emanuel predicted that autonomous AI would probably exceed physicians and physician–AI teams across five cognitive medical functions and might be ready for some or many workflows by 2030. Simulated interviews, difficult diagnostic cases, test-selection exercises and treatment studies are assembled into a forecast about the best medical care; adverse findings often lose force because their models were older or less elaborately configured, while favourable findings retain their standing as evidence of what comes next. The future is unusually cooperative with the conclusion. The underlying studies also require a more immediate inquiry into what the apparently autonomous system contains.5
One chronic-disease example makes the methods section worth the interruption. In the 32-person insulin-titration trial, an Alexa-based interface delivered instructions from a deterministic, rules-based dosing system; the clinician had selected the protocol before activation, specifying the starting dose, glucose target and adjustment instructions. Frequent interaction helped put that policy into effect, and titration improved relative to usual care. The clinician’s judgment was already in place before the machine began its conversations; she can scarcely be required to read the instructions aloud each morning for that earlier judgment to remain part of the treatment. Whether the encounter now sounds autonomous tells us very little about where its clinical decisions were made.6
The clinician can disappear from the other end of the process as well. Guidance from the International Council for Harmonisation asks trialists to specify the treatment effect they intend to estimate—the estimand—including how rescue treatment and changes of therapy enter the assessment. A service tested with human rescue available answers a different question from the hypothetical service without it. Applied to clinical AI, the distinction permits a successful rescue to count towards the benefit of a deployed service while requiring us to retain the people who performed it among the service’s components. A clinician who prevents an error from reaching a patient has done work; crediting the outcome to the autonomous model is an unusually efficient way to make that work disappear.7
We use construct capture for the arrangement in which an interested party controls enough of the task definition, sampling, grading and interpretation that success under its instrument becomes evidence of a broader capacity without an independent validity argument. An excellent research assistant can win an excellent test of research assistance; describing its winner as the leading clinical model extends the finding to work the test may scarcely contain. The description grows: the system performs strongly on these tasks; it is good at clinical work; it possesses clinical intelligence; it is safe enough to decide. As the sentences grow, the score stays still. Its jurisdiction does not.
II. Who wrote the right answer?
Suppose we make the claim more modest: this benchmark assesses whether a system follows a named clinical guideline. The source is now identified, and the evaluator still has to decide what it requires in each case—how its conditions, exceptions and degrees of recommendation apply, and where it returns a decision to judgment. A benchmark is itself a compiler: it translates a conception of competence into tasks and criteria, just as making the clinical reference executable translates prose into rules. Both translations contain decisions. The benchmark can test the model’s reading only after its authors have supplied a reading of their own.
Consider an invented instruction of a familiar kind: Consider specialist assessment if symptoms persist after adequate treatment. A patient has continuing symptoms, and the model says she must be referred. The word referral appears, to the satisfaction of an answer key looking for it, although we have yet to establish what counts as persistent, which treatment would have been adequate, whether the patient received it or what happened to the choice preserved by consider. The model and the grader can agree on ‘refer’ while sharing the same unsupported premises. Their agreement, reassuring as it looks, may be the very thing we need to investigate.
Some of the missing work is ordinarily done by the reader. Piantadosi, Tily and Gibson explain how ambiguity can make communication efficient where context already supplies information: repeating everything in every sentence would be wasteful, and the listener’s contribution permits the writer to economise. A clinician may know where the duration is defined, which treatment the preceding passage discusses and what local service can provide the assessment. Giving the sentence to a model changes who must supply that context. It may recover the intended meaning from the document, infer a plausible meaning from elsewhere or import a familiar rule from another jurisdiction, with equally assured prose accompanying each. Somewhere in this operation, plausibility, intention and local authority can become difficult to distinguish; the record needs to tell us which has actually been established.8
Roman Jakobson gives the translator a concrete version of the difficulty. ‘I hired a worker’ leaves information unspecified that its Russian translation, in his example, requires about gender and aspect; the receiving language asks questions the original could leave unanswered. Our software asks questions of its own: a field expects a duration, a condition a truth value, an action a defined force. The guideline may have supplied none of these in the required form. We can preserve the unresolved value, obtain an approved interpretation or insert something that makes the program run. In the last case, a required field has become an amendment to the guideline, with the considerable convenience that the people entitled to approve an amendment may never see it.9
The question addressed to the system can contribute an unsupported premise just as readily. ‘What should we do after her failed course of treatment?’ invites the respondent to consider the next intervention before establishing what made the treatment a failure. Linguists call one mechanism at work here presupposition accommodation: information presented as already shared is accepted in the effort to understand what is being proposed. The same facility serves ‘Now that AI has reached clinical competence, how quickly should hospitals deploy it?’, which offers the timetable for discussion with competence already entered as an agreed fact. The reader who wants evidence must first interrupt the sentence’s proposed business. A clinical assessment ought to reward that interruption when the premise matters, even if the response is less immediately obliging than the questioner had hoped.10
The remedy depends on the kind of gap. A missing patient fact may call for a question or chart search; an unclear guideline term may call for another passage or a clarification by those responsible for adopting the policy; deliberate discretion needs to remain available to the person who holds it. In our example, fixing the local meaning of adequate treatment is a policy decision, establishing whether the patient received it a factual inquiry, and considering referral the subsequent clinical act. Keeping those steps apart gives the reviewer something specific to approve or challenge, and prevents a fluent answer from making all three decisions at once without showing where they occurred.
The resulting clinical specification is an explicit account of the patients to whom a standard applies, the facts needed to use it and the actions it requires, permits, prohibits or leaves to judgment. It becomes executable when software can apply those conditions and record the rule used. Attached should be the source passage, each consequential interpretation, its approving person or body and the version in force. A benchmark built from this reference would expose why it expects the answer, including the decisions its authors made in places where the published guideline remained open.
A scholarly edition supplies a surprisingly exact model for retaining those decisions. An editor prints a reading and records the alternatives, the manuscripts supporting them and the people who proposed a change. In the Text Encoding Initiative’s worked example from Propertius, Stephen Heyworth leaves two lines where they are while recording A. E. Housman’s proposal to move them. The page commits; the apparatus preserves the disagreement. Housman may well have been right about where the lines belonged; the apparatus is hospitable enough to let the reader reconsider the matter. A clinical counterpart—an executable critical edition—would give the clinician inheriting a local interpretation a comparable opportunity to discover what has been moved, by whom, and why, with the hospital’s amendment identified as such, its approval attached and its departure from the source available for inspection.11
An unresolved interpretation may still permit useful action. Work on underspecified semantics, including Koller and Thater’s weakest readings, examines how several readings can remain represented while conclusions supported in common are recovered. The clinical application requires complete obligations, including timing, discussion and monitoring. Where every retained reading permits the same next step on the known facts, that step may be supported while another question stays open. Where the readings require different conduct, the system can produce the patient circumstances on which they diverge and submit those circumstances with the relevant passages for review. The reviewer receives an answerable question; she must also judge whether the alternatives are adequate, since listing two does little good if the source admits a consequential third. We would like the compiler to preserve that inconvenience. It belongs to the medicine we asked it to preserve.12
This gives the answer key a basis that can be examined separately from the model’s response. It also gives us, having proposed to build such a thing, a burden of our own: the formalisation must preserve the source, and the implementation must preserve the formalisation. Precision makes these questions easier to ask. It can also make an interpretive error exceptionally faithful to itself.
III. What checking can establish
Once the rules are explicit, we can ask mathematical questions of them: whether a permitted combination of facts produces contradictory instructions, whether the declared population contains possible patients and whether the circumstances meant to trigger a branch can actually reach it. These checks concern what we have represented. Fidelity to the source requires another inquiry—whether the duration is the intended one, the exception survived and the discretion in consider remains available. A program can satisfy the mathematical checks while implementing the wrong reading perfectly, so the word verified needs to be accompanied by an account of what was verified.
DeepSeek’s theorem-proving work makes the point with unusual economy. Twelve reported benchmark successes became ten after two formal problem statements proved defective. In a published example, a condition intended to specify a sequence’s minimum positive period also prohibited a shift of zero, which leaves every sequence unchanged; the model recognised the contradiction and completed a valid proof from it. The model had done something intelligent and the proof machinery something correct, while the advertised achievement still needed revision because the assumptions described an impossible object. In clinical software, a corresponding defect might prevent any patient from satisfying the referral condition. A guarantee covering everyone who satisfies it would then cover nobody. Establishing that the population is satisfiable and its intended branches reachable gives the vocabulary an immediate practical purpose: we should like the patients protected by the guarantee to be capable of existing.13
Source fidelity presents another difficulty, illustrated by Catala, a programming language for legal rules. Its authors encoded French family benefits and initially agreed with a state-sponsored simulator on their test cases; examining the underlying code then exposed an omitted income-cap exception for one-child families in overseas territories, which was reported and corrected. The exception was already in the legislation. Had the earlier simulator supplied a benchmark’s answers, a successor faithful to the source could have lost points for disagreeing with it. Clinical systems can inherit a mistaken interpretation in the same way, particularly when agreement among existing implementations supplies the next answer key: the software defect survives, with the authority of the guideline still attached.14
Source review, logical checking and implementation testing therefore need distinct evidence. Clinical reviewers examine whether the adopted reading is justified, formal checks examine its represented rules, and tests of the final system examine what survives execution, extraction of patient information and presentation of the result. Independence matters here, including independence in the ways things can go wrong. In a 1986 experiment, Knight and Leveson found that failures among twenty-seven independently developed versions of a program coincided more often than statistical independence predicted. Separate authorship had left room for shared mistakes. Three translators working from the same faulty edition may agree with considerable confidence on something the author never wrote; commissioning a fourth can make the consensus more impressive while leaving the edition exactly as it was. The useful question is what, independently, the next reviewer has checked.15
We can put some deliberate pressure on the checking. Translation validation offers a precedent for examining whether each generated program preserves a specified semantic relation to its source, once that source has been made formal; the interpretation of the medical prose remains the prior decision. Then change something on purpose. Alter the duration, remove an exception, replace permission with obligation, and see whether the assessment objects. In the referral example, turning consider assessment into always refer should have consequences if the adopted policy preserves a choice. An assessor that catches the planted defect supplies evidence of its ability to detect that failure; an assessor that endorses every version has told us something useful too, although its maker may find the result less agreeable.16
Medicine already has an institutional programme for this work. NICE distinguishes offer, consider and discuss; WHO’s SMART Guidelines separates narrative guidance from operational and executable forms; HL7 places intermediate representations before clinical and technical reviewers. The Internet Engineering Task Force has undertaken its own negotiations over normative language: RFC 2119 defines MUST, SHOULD and MAY, and RFC 8174, twenty years later, clarifies that those meanings apply when the words are capitalised. The internet needed another standards document to settle what its standards documents meant by should. Each of these efforts makes a transformation available for inspection before the next one obscures it, doing in stages what a conversational interface can appear to accomplish in a single response. The apparent category is software that reads medical text; the deeper work is specification recovery, semantic compilation, completeness and governance beneath the linguistic surface.17
The discussion of ‘neuralese’ brings us to the other side of this translation. One recent argument invoked Magnus Carlsen, whose recognition of a position can exceed his ability to narrate the process by which a move became apparent. Grant the point. His opponent can still see the board, the permitted movements of the pieces have been settled, and the move itself remains available for examination however privately it was conceived. The arbiter can establish that a move is legal while having very little idea why anyone would want to play it. An account of the player’s inspiration and an account of what he was entitled to do answer different questions.18
Latent-reasoning research gives this distinction a technical setting. A model can feed an internal representation back into further computation before producing another word, or repeatedly apply a computational block so that additional work occurs between the verbal steps available to an observer. ‘Neuralese’ is a loose name for this possibility, carrying rather more suggestion of an undiscovered language than the evidence warrants. The consequential fact is that the work can increasingly proceed beyond the account we are able to read. The paragraph on the screen becomes one representation whose relationship to that work needs to be established.19
Astra’s system card makes the oversight problem concrete: OpenAI reports reduced chain-of-thought monitorability and greater control over what the model reveals, including experiments in which it was prompted to conceal deliberate underperformance. Access to the actions themselves improved detection in some settings. The record offered for inspection can therefore become another output the system knows how to manage; a sufficiently accomplished explanation may require an explanation of its own.20
An account may be intelligible, faithful to the computation, useful to a monitor or sufficient to establish clinical warrant, with separate evidence required for each. Return to our referral: whether the patient received adequate treatment, whether symptoms have persisted for the relevant interval and whether consider has been preserved are questions that remain answerable against an adopted specification, however the model reached its proposal. We should like the system to reason as well as it can. We should also like the obligations it must satisfy to remain available when its reasoning becomes harder to follow.
The rule can be enforced through different architectures. One system might execute approved rules directly; another might allow an AI agent to propose actions while a separate mechanism checks constraints and provides a safe fallback. NASA’s formal runtime-assurance work on automated braking shows how monitors and fallback behaviour can support safety claims under explicit assumptions. The clinical implementation owes us evidence for its own assumptions: that it identifies the relevant patient facts, intervenes in time and has somewhere safe to send the case. The benchmark should assess these obligations whichever architecture is used. Deterministic execution of a defective rule deserves to fail alongside an eloquent model that ignores a correct one.21
IV. From a correct answer to care
A specified system following an approved interpretation under tested conditions has demonstrated conformance. What that achievement does for a patient depends on how the system obtains the facts and helps bring about the action. Our referral rule may be impeccably represented, yet a chart recording what was prescribed supplies only part of the history needed to establish treatment failure. Whether the patient received the treatment, and whether the information available supports that conclusion, remain questions the quality of the rule cannot answer on her behalf.
MT-InfoSeek names the relevant informational achievement final sufficiency: whether what the model has learned determines the target independently of the answer it eventually produces. In clinical assessment, a missing fact matters when its possible values would change what the standard supports. This gives questioning a purpose and a stopping point, taking account of the burden and delay of further inquiry. We have no wish to reward a system for subjecting the patient to an exhaustive interview merely because exhaustiveness is easy to score. The system needs enough information to warrant what it proposes; one that recommends referral by guessing the missing premise can match the answer key and still have failed to establish that it knew enough to recommend it.22
Even adequate facts and faithful execution leave the combined plan to be examined. In 2005, Cynthia Boyd and colleagues applied disease-specific guidelines to a hypothetical 79-year-old woman with five chronic conditions and obtained twelve medications, a demanding regimen and opportunities for adverse interactions. The paper concerned quality measurement and pay for performance, which gives our new benchmark problem a somewhat older address. Five modules can each execute as intended while the patient must live with the combined result, including its burden, competing objectives and demands upon her priorities. A benchmark covering one guideline can make a useful claim within that scope; one advertised as an assessment of integrated management owes us tests of what happens when the recommendations meet in the same patient.23
Scope also determines how far the assessment follows an action. A tool promising referral advice can be assessed on its advice; a service promising to manage referrals has undertaken to get the request to someone, observe an appropriate time limit and do something when the service fails to materialise. Software verification distinguishes constraints on behaviour from requirements for eventual progress, the latter called liveness. A system can respect every prohibition while returning ‘refer for review’ indefinitely, leaving the patient in a state of impeccably documented neglect. A correct deferral creates work that somebody must receive, own and be in a position to complete.24
The record should preserve how that work actually unfolded. Luciana Duranti’s archival distinction between authenticity—a record being what it purports to be—and reliability reminds us that an authentic record can faithfully preserve an unsupported assertion. Computational reproducibility introduces another question. Monday’s account needs the lab result available on Monday, the rule then in force and the action taken; Friday’s correction belongs in the record as a later event. A rerun with Friday’s knowledge may produce a better recommendation, and we should want it to, while retaining the evidence needed to explain Monday’s decision. The patient attended on Monday. Giving her clinician the benefit of information available four days later would improve the apparent quality of the earlier consultation at the expense of recording what occurred.25
These distinctions explain how a collection of benchmark victories can leave patient benefit unresolved. Medicine encounters a related difficulty with surrogate endpoints: a marker can be measured accurately while its relationship to how a patient feels, functions or survives remains incompletely established. The analogy is limited but useful. A benchmark offered as evidence of better care needs a demonstrated connection to care in the proposed setting, to which source fidelity, conformance, communication and workflow evidence can make distinct contributions. A study of the completed service must also keep human rescue and downstream work attached to the result according to the question it evaluates. However many additional scores we obtain, a relationship their experiments never examined remains a relationship requiring investigation.26
FDA’s August discussion paper on generative-AI-enabled devices asks questions organised around that use. It considers benchmarking and clinical confirmation of the final configured device, proportionate to purpose and risk, and asks whether the relevant comparator is a clinician, a clinician using the device or an autonomous system. These are different experiments because they are different products in different distributions of authority. The paper also asks what establishes construct validity and how sponsor-developed tests, sequestered cases and independent adjudication should be treated. Its questions remain open. Read beside the Emanuel Perspective, the documents undertake different kinds of preparation: one anticipates the organisation of care after the evidentiary gaps have been closed; the other asks what would close them.27
V. Who controls the standard?
A hospital presented with this evidence should be able to identify the system and use tested, inspect the clinical interpretation behind the answers, and distinguish a defect in that interpretation from an error in execution or a failure to deliver the care. Even then, someone must decide whether the evidence is sufficient for the proposed reliance. And when experience exposes a mistake, someone must be able to correct the standard against which everyone has been performing so successfully. Both are institutional acts on which the score’s practical importance depends.
An examination has long occupied a bounded place in professional permission. In Dent v. West Virginia, the Supreme Court upheld a scheme in which education, experience or examination could establish qualification, while the state’s certificate conferred authority to practise. Evidence of competence and the institution empowered to act on it belonged to the same process, with different jobs. Authority is a proposed transfer of power: who may act, within which scope, under which duties and before which body the actor must answer. There is considerable institutional work inside a medical licence, however efficiently a model completes the questions; for clinical AI, the corresponding inquiry concerns the use institutions will permit and the responsibilities incurred by permitting it.26
A benchmark can obtain practical authority through other people’s reliance on it. Once passing becomes a condition of purchase or deployment, control over the answer key helps determine which products can be accepted, a possibility made unusually concrete in Allied Tube v. Indian Head. Steel-conduit interests recruited 230 people into the association whose electrical code was routinely adopted by governments, then coordinated their votes against plastic conduit through walkie-talkies and hand signals; the delegation included a national sales director’s wife, apparently summoned by the urgency of electrical safety. The proposal failed, 394–390. The jury found compliance with the association’s literal rules, a genuine safety motivation in part and subversion of the consensus process—findings worth reading together—and the Supreme Court rejected the claimed antitrust immunity. What happened in the meeting mattered because others relied on its outcome, which is precisely the standing a benchmark seeks when its result becomes a condition of use.28
The clinical consequence concerns who can challenge that standing. A benchmark owner needs the ability to repair a faulty interpretation and an obligation to answer when a competitor, clinician or patient disputes it. Expertise, commercial interest and sincere concern for safety can coexist quite comfortably; the procedure must work under those ordinary conditions, including when the person best placed to identify a defect stands to benefit from its correction. Construct capture has reached its second stage when control over measuring a capacity becomes control over what conduct institutions will accept as evidence of possessing it.
Effective challenge begins with access to the rule. In Public.Resource.Org v. Commission, the Court of Justice of the European Union required disclosure of requested harmonised standards because their legal effects created an overriding public interest. The holding belongs to its legal setting; the broader governance question is how anyone can challenge a consequential specification they are unable to inspect. A clinical benchmark can publish the obligations it tests while keeping its examination cases unseen. Preventing a developer from memorising the next patient is compatible with letting her know what the system will be required to do for that patient. The confidentiality needed to protect an examination should have a defined object and a reason.29
The rule also needs a future beyond its first implementation. The OECD’s Law-as-Code consultation asks how authorised executable legal rules could become shared infrastructure while keeping interpretation, amendment and maintenance visible. Our clinical proposal raises the corresponding questions: who approves a local variant, who can restore an omitted exception, what becomes of the standard if the vendor withdraws? A company may reasonably charge for doing this work well (a proposition with which this author is entirely comfortable). The institution relying on the result still needs to inspect its commitments and arrange their continued stewardship. A commercial relationship can end while the hospital continues to have patients and software applying rules in its name.30
FDA’s Accreditation Scheme for Conformity Assessment supplies an existing arrangement for private evidence under defined public oversight: recognised standards, competent laboratories, testing scopes and controlled reports, with accreditation withdrawn when its conditions fail. Withdrawals following audits in 2025 included cases involving data-integrity concerns. Independence, here, is a scoped, audited and revocable institutional status. A safety assurance case supplies the complementary argument: the claim being made, its supporting evidence, the assumptions on which it depends and the objections that could defeat it. Benchmark findings can occupy specified places in that case alongside evidence about the clinical source, implementation and delivered service. The rest of the argument remains somebody’s responsibility.31
Having warned that private tests can become public standards, we now introduce a private test and hope it matters. GuideBench, which we are developing at Compiled Health, is intended to assess whether a clinical system obtains the information needed for a decision and satisfies a named, approved clinical specification. We build executable representations of guidelines, so this is, unsurprisingly, a conception in which our work looks useful. It also gives us the opportunity to reproduce an error with great consistency if our implementation and our answer key inherit the same mistaken reading. Source fidelity, logical and implementation soundness, and the tested system’s conformance must remain separately examinable, especially when they happen to agree.
For the patient whose symptoms persist, assessment begins with the adopted meaning of the referral rule. Cases can vary symptom duration, treatment received and which relevant facts are available; a contestant can ask in a different order or express its conclusion in different prose, provided it uses the information it has, obtains what remains necessary, preserves the choice in consider and meets the further obligations included in its declared workflow. An unsupported assumption that treatment failed should be distinguishable from misreading the rule, and a justified request for information from a lucky answer. The point is to locate the success or error precisely enough that we know what the result establishes—and, where necessary, what to repair.
The reference would retain its source, adopted reading, unresolved alternatives, local amendments and approval history: the apparatus of an executable critical edition. Challenges could reach the interpretation as well as the marking, and a material dispute could suspend an affected result pending adjudication. Assessors would need sufficient independence to expose shared failure mechanisms; changes to a guideline, model, surrounding software or patient-data mapping would have defined consequences for reuse of the evidence. Our preferred architecture would face the same requirements as a competitor’s. We could announce our impartiality, of course, although a competitor entitled to make us correct the reference would be a more persuasive witness.
The resulting report would identify the system, interpretation, cases and patient information, and state where the system succeeded, failed or needed more information. Ordinary-care results, safety-critical failures and frontier challenges would remain distinct, with overall rank omitted by design. A narrow system could conform throughout its scope while a more capable model departed from the policy; the departure could then be examined for its clinical merit with the conformance finding intact. Evidence of benefit and permission to deploy would retain their places in the larger assurance process. Another institution would receive a claim it could examine and decide how to use.
At the bedside, the software would still need to record the interpretation and facts actually used, so that a clinician could locate a misread fact, challenge a rule or depart from it on patient-specific grounds. The benchmark’s report must remain available for the same reason: somebody may have cause to ask what, exactly, was established, including someone whose patient has given her cause to doubt it. A benchmark can begin by describing competence and end by helping to govern it. Once it proposes to speak in the profession’s name, the profession should be able to answer back—and to make the correction count.
Notes
- 1OpenEvidence, ‘OpenEvidence Creates the First AI in History to Score a Perfect 100% on the United States Medical Licensing Examination’, company release, 15 August 2025; Rebecca Soskin Hicks et al., ‘HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats’, arXiv:2604.27470v1, 30 April 2026; Louis-Antoine Mullie, ‘Which AI Gives the Safest Medical Advice?’, Doximity, 17 July 2026. These establish reported results under different protocols, including different system configurations.↩
- 2Hicks et al., ‘HealthBench Professional’, §§3–5. The original report distinguishes base GPT-5.4 from ChatGPT for Clinicians using GPT-5.4 and describes weighted criterion scoring and length adjustment.↩
- 3American Educational Research Association, American Psychological Association and National Council on Measurement in Education, Standards for Educational and Psychological Testing (2014), ch. 1; Michael T. Kane, ‘Validating the Interpretations and Uses of Test Scores’, Journal of Educational Measurement 50 (2013): 1–73, doi:10.1111/jedm.12000. The clinical examples apply their distinctions.↩
- 4David Wu et al., ‘First, do NOHARM: A Medical Safety Benchmark and Randomized Study of Physician and AI Teaming on Clinical Consultations’, arXiv:2512.01241v4, 13 July 2026. The study reports omissions as more than 80 per cent of severe errors. These assessments concern potential harms of applying recommendations.↩
- 5Ezekiel J. Emanuel, Abe Baker-Butler, Neal Khosla and Vinod Khosla, ‘Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care?’, JAMA, published online 17 August 2026, doi:10.1001/jama.2026.15380. The article is a Perspective. Its five-task synthesis, handling of contrary evidence, caveats and concluding forecast are the objects of the criticism. The authors’ affiliations include Curai Health and Khosla Ventures.↩
- 6Ashwin Nayak et al., ‘Use of Voice-Based Conversational Artificial Intelligence for Basal Insulin Prescription Management Among Patients With Type 2 Diabetes: A Randomized Clinical Trial’, JAMA Network Open 6 (2023): e2340232, doi:10.1001/jamanetworkopen.2023.40232, Methods, ‘Voice-Based Conversational AI’. The trial specifies a deterministic, rules-based dosing system and clinician selection of the protocol before activation.↩
- 7International Council for Harmonisation, E9(R1): Addendum on Estimands and Sensitivity Analysis in Clinical Trials, final version adopted 20 November 2019, §§A.3.1–A.3.3. Applying its distinction among treatment-policy, hypothetical and other estimands to AI oversight and rescue is the proposal made here.↩
- 8Steven T. Piantadosi, Harry Tily and Edward Gibson, ‘The Communicative Function of Ambiguity in Language’, Cognition 122 (2012): 280–91, doi:10.1016/j.cognition.2011.10.004. The argument concerns communication with informative context. Its application to clinical compilation is proposed here.↩
- 9Roman Jakobson, ‘On Linguistic Aspects of Translation’, in Reuben A. Brower, ed., On Translation (Harvard University Press, 1959), 232–39. The comparison between obligatory grammatical distinctions and software schema requirements is an analogy developed in this essay. The invented referral instruction illustrates the interpretive problem.↩
- 10David Lewis, ‘Scorekeeping in a Language Game’, Journal of Philosophical Logic 8 (1979): 339–59; Kai von Fintel, ‘What Is Presupposition Accommodation?’, MIT manuscript, 2000. The deployment and clinical questions are invented illustrations.↩
- 11Text Encoding Initiative, TEI P5 Guidelines, ch. 13, ‘Critical Apparatus’, transposition example at Propertius 1.16, showing Stephen Heyworth’s treatment of A. E. Housman’s proposed relocation of two lines. The clinical counterpart is our proposal.↩
- 12Alexander Koller and Stefan Thater, ‘Computing Weakest Readings’, Proceedings of ACL 2010, 30–39. Their technical result concerns scope-underspecified representations and entailment. Extending the approach to clinical obligations requires explicit treatment of modality, admissible interpretations, patient states and accompanying duties.↩
- 13Z.Z. Ren et al., ‘DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition’, arXiv:2504.21801v2, §3.3 and Appendix C. The revised paper reports the correction from twelve to ten CombiBench problems and illustrates proof from contradictory hypotheses, including the minimum-period example. The clinical analogy concerns satisfiability and reachability within the represented specification.↩
- 14Denis Merigoux, Nicolas Chataing and Jonathan Protzenko, ‘Catala: A Programming Language for the Law’, arXiv:2103.03198v2, §6. The benchmark derived from the earlier simulator is a counterfactual illustration.↩
- 15John C. Knight and Nancy G. Leveson, ‘An Experimental Evaluation of the Assumption of Independence in Multi-Version Programming’, IEEE Transactions on Software Engineering SE-12 (1986): 96–109, doi:10.1109/TSE.1986.6312924. The experiment examined twenty-seven independently developed versions using a million tests. The extension to assessor design concerns the assumption of independent error; the historical failure rates are not assigned to contemporary AI. The translators are an invented illustration of a shared source of error.↩
- 16Amir Pnueli, Michael Siegel and Eli Singerman, ‘Translation Validation’, Tools and Algorithms for the Construction and Analysis of Systems, LNCS 1384 (1998), 151–66. The method checks an individual translation under a defined semantic relation. Its source is a formal program; the prior adoption of an interpretation of medical prose is a distinct step. The proposed duration and exception mutations are clinical applications of adversarial testing.↩
- 17Scott Bradner, RFC 2119 (1997); Barry Leiba, RFC 8174 (2017); NICE, Developing NICE Guidelines: The Manual, ‘Interpreting the Evidence and Writing the Guideline’; WHO, ‘SMART Guidelines’; HL7 International, Clinical Practice Guidelines Implementation Guide, version 2.0.0, ‘Approach’.↩
- 18Scott Stevenson, post on X concerning Magnus Carlsen, intuition and human legibility, reproduced in the research materials for this essay; FIDE, Laws of Chess, effective 1 January 2023, Preface and Articles 3, 8 and 12. The opening grants the post’s premise for the argument; it offers no independent study of Carlsen’s cognition. The comparison concerns the external specification and adjudication of a move, whose legality can be examined separately from its strategic merit. FIDE’s preface, incidentally, reserves room for an arbiter’s judgment where the laws leave a situation unsettled. Even chess has remembered to say who should decide what its rules have left open.↩
- 19Shibo Hao et al., ‘Training Large Language Models to Reason in a Continuous Latent Space’, arXiv:2412.06769, introduced in 2024 and accepted to COLM 2025; Jonas Geiping et al., ‘Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach’, arXiv:2502.05171 (2025). Coconut feeds hidden states back as inputs in continuous space; the recurrent-depth model iterates a computational block at inference time. These papers establish examples of latent computation. Identifying Astra’s architecture or establishing a private language with a vocabulary and grammar would require separate evidence. The implications for clinical conformance are developed in the essay.↩
- 20OpenAI, GPT-6 Astra System Card, September 2026, §9, especially §§9.1–9.3. The report distinguishes chain-of-thought-only, action-only and full-context monitoring, including prompted evasion experiments. The detection gains from additional context depend on the task and monitor; the card also reports settings where covert underperformance was difficult to detect. These evaluations supply evidence about the tested systems and conditions. The interpretation that a verbal trace can itself become an output managed by the model is the argument developed in the body, as is the distinction between such evidence and clinical warrant.↩
- 21J. Tanner Slagel et al., ‘A Formal Verification Framework for Runtime Assurance’, NASA Technical Reports Server, document 20230017350; NASA Formal Methods Symposium, 2024. The work formalises Simplex runtime assurance and illustrates it with automatic vehicle braking. The clinical analogue is conditional on its own state-estimation, timing, enforcement and fallback assumptions.↩
- 22Yepeng Huang et al., ‘Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking’, arXiv:2608.14808, 14 August 2026. Final sufficiency is assessed independently of the answer generated. Source uncertainty, patient burden and recommendation-specific warrant are proposed extensions.↩
- 23Cynthia M. Boyd et al., ‘Clinical Practice Guidelines and Quality of Care for Older Patients With Multiple Comorbid Diseases: Implications for Pay for Performance’, JAMA 294 (2005): 716–24, doi:10.1001/jama.294.6.716. This was a hypothetical patient and historical guideline analysis; the inference about composition of current conformance assessments is ours.↩
- 24Bowen Alpern and Fred B. Schneider, ‘Defining Liveness’, Information Processing Letters 21 (1985): 181–85, doi:10.1016/0020-0190(85)90056-0. Eventual progress is a liveness property; a concrete deadline can be expressed as a safety property because its violation becomes observable in finite time. Clinical escalation ownership, time limits and completion checks are proposed workflow requirements.↩
- 25Luciana Duranti, ‘Reliability and Authenticity: The Concepts and Their Implications’, Archivaria 39 (1995): 5–10; Christian Chartier, ‘Reasoning Is Not Language’, Compiled Health, 19 July 2026. Historical authenticity, computational reproducibility and clinical justification are distinguished here. The Monday–Friday illustration is invented.↩
- 26FDA, ‘Surrogate Endpoint Resources for Drug and Biologic Development’; Dent v. West Virginia, 129 U.S. 114 (1889). The surrogate-endpoint comparison is an analogy; the FDA doctrine concerns drug and biologic development. The legal example concerns the institutional use of qualification evidence.↩↩
- 27FDA, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback, August 2026, especially §V and questions 10–16. The paper invites discussion and establishes no new binding requirements.↩
- 28Allied Tube & Conduit Corp. v. Indian Head, Inc., 486 U.S. 492 (1988), 496–98, 509–11. The analogy concerns private standard-setting power and its governance.↩
- 29Public.Resource.Org, Inc. and Right to Know CLG v. European Commission, C-588/21 P, judgment of 5 March 2024, ECLI:EU:C:2024:201, especially paragraphs 65–89. The Court found an overriding public interest in disclosure of the requested harmonised standards owing to their legal effects. The broader question of standards copyright remains distinct from this holding. The clinical publication requirement proposed here is a governance argument.↩
- 30OECD, ‘Consultation on the Digital Provision of Law: Towards a Shared Reference Framework for Law as Code’, opened 29 July 2026. This is a consultation.↩
- 31FDA, ‘Accreditation Scheme for Conformity Assessment’; NASA, ‘New Tool for Developing Safety Assurance Cases’, 22 September 2020; AAMI TIR38:2019, Medical Device Safety Assurance Case Guidance, FDA recognition 6-422. These concern defined testing scopes and structured assurance.↩