Commitment two, in the blueprint, is nine words:
Scepticism is trained explicitly, and it is assessed. Not a lecture on limitations — a drilled reflex, like recognising a deteriorating patient.
Of the ten commitments this is the one I have leaned on hardest and defended least. Every other companion post in this series has been about machinery — how you set a pass mark, how you build an examination, how you measure whether any of it survived contact with a ward. This one is about the thing the machinery is for. It is also the only commitment that the trial I keep citing directly implicates, because that trial is a study of people who had been taught about AI and were harmed by it anyway.
So this post does four things. It reads that trial properly, including the parts that weaken it. It assembles the wider evidence, which is older and broader than anyone writing an AI curriculum in 2026 seems to acknowledge. It then confronts a literature that ought to terrify everyone in this field and is almost never cited in it — the medical education research showing that taught debiasing does not work. And it sets out what I would build instead, in enough operational detail that somebody could disagree with a specific number rather than with a mood.
I will say the conclusion at the top, because it is the part I would most like argued with.
Part 1 — The trial, read properly
The study is Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy, by Qazi and colleagues, published in NEJM AI in 2025 and available as a preprint. It was prospectively registered (NCT06963957) before enrolment, with the statistical analysis plan uploaded in advance. I want to set out what it actually did, because I have seen it cited three times now as though it showed that AI makes doctors worse, which is not what it showed.
The population. Forty-four physicians registered with the Pakistan Medical and Dental Council, recruited from two consecutive cohorts of a training programme at the Lahore University of Management Sciences. Median clinical experience ten years. Every one of them had completed a twenty-hour AI-literacy course covering large language model capabilities, prompt engineering, and strategies for critically evaluating AI-generated output. This is the feature that makes the trial worth building a curriculum around: the participants were not naive. They were the product of exactly the intervention most health systems are currently commissioning.
The design. Randomised 1:1, single-blind, participants blinded to the study's aims. Six clinical vignettes across internal medicine, cardiology, neurology, paediatrics, infectious disease and emergency medicine, developed by three physician co-authors, with the too-easy and the too-rare deliberately excluded. Seventy-five minutes. The control arm received unmodified ChatGPT-4o output on all six cases. The treatment arm received output into which subtle but clinically significant errors had been embedded on three of the six — errors designed to be detectable by a competent physician but not obvious on casual reading — with the position of the erroneous cases randomised so that no pattern could be learned.
Three features of the design matter more than the headline number.
First, consultation was voluntary and required an explicit click. Nobody was forced to look at the model. This is the on-demand pattern that is actually spreading in clinical practice, as opposed to the mandatory-review or always-visible designs used in most earlier work, and it is the pattern that preserves the autonomy everyone invokes when they say the clinician remains responsible.
Second, both arms kept their ordinary resources — medical databases, standard search — with a browser extension specifically blocking Google's AI Overviews so that the control arm's exposure was genuinely controlled. That is a level of methodological care I would not have expected.
Third, the primary endpoint was reasoning, not the answer. Participants documented their top three differentials with supporting and opposing evidence, their top choice with justification, and recommended next steps. Three blinded physicians scored each response against a rubric developed by independent solution of every case and consensus resolution of discrepancies. Inter-rater reliability was high (Krippendorff's α = 0.93) and the instrument's internal consistency was strong (Cronbach's α = 0.80).
The result. Mean diagnostic reasoning accuracy was 84.9% in the control arm and 73.3% in the treatment arm: an adjusted difference of −14.0 percentage points (95% CI −18.9 to −9.1; P < .0001) from a pre-specified linear mixed-effects model with random effects for participant and for case. On the secondary endpoint, top-choice diagnostic accuracy, 90.5% versus 76.1%, an adjusted −18.3 percentage points (95% CI −26.6 to −10.0; P < .0001).
Now the detail that I think is the most important single number in the paper, and which almost nobody quotes.
The subgroups, which are where the paper gets interesting and where I would be most careful. Physicians at or above the median of ten years' practice were harmed more (−16.6 percentage points, 95% CI −23.1 to −10.1) than those below it (−9.1, 95% CI −18.1 to −0.1). Physicians using large language models at least weekly showed a significant decrement (−11.0, 95% CI −18.5 to −3.6); infrequent users' point estimate was almost identical (−10.7) but its interval crossed zero (95% CI −24.5 to 3.1), which is a statement about sample size rather than about infrequent users. And male physicians showed −25.8 (95% CI −33.8 to −17.7) against −2.1 in female physicians (95% CI −9.8 to 5.5, not significant).
What I would and would not teach from this
The honest summary of the trial's limits, most of which the authors state themselves:
- Forty-four people. The recruitment target was fifty; the observed effect was large enough that the study remained adequately powered, but this is a small trial and the confidence intervals are correspondingly generous.
- Vignettes, not patients. No time pressure of the real kind, no multimorbidity, no relatives in the corridor, no consequence for being wrong.
- The errors were ours, not the model's. Errors introduced deliberately by a panel of physicians are the errors physicians can imagine. Real model failures may be more subtle, or differently shaped, or concentrated in places we would not think to seed.
- One model, one session. ChatGPT-4o, in mid-2025, in a single 75-minute sitting. Whether the effect grows or shrinks with months of use is exactly the question the design cannot answer.
- Physicians only. No nurses, no clinical officers, no pharmacists — which for an institution whose largest and highest-yield cadre is nursing is a substantial gap.
- No mitigation was tested. The trial establishes the problem. It says nothing whatever about the fix, which is the entire subject of this post.
What survives all of that is the direction and the rough order of magnitude. A pre-registered randomised design, blinded graders, high inter-rater reliability, a pre-specified primary endpoint, an effect of fourteen points with an interval nowhere near zero, and an equal consultation rate that closes off the easy explanation. I am comfortable building on the claim that AI-literacy training did not protect these physicians. I am not comfortable building on any number in the paper to two significant figures, and the curriculum should not depend on one.
Part 2 — This is not one study
The reason I am willing to design an institution around a 44-person trial is that it is not load-bearing on its own. It is the most recent and most directly relevant point on a line that runs back through radiology, dermatology, pathology and endoscopy to thirty years of aviation human factors. The consistency is the argument.
The taxonomy, from the cockpit
The vocabulary comes from Linda Skitka, Kathleen Mosier and colleagues, working on glass-cockpit aviation in the late 1990s. Their Does automation bias decision-making? and Accountability and automation bias established the split that still organises the field:
- Errors of omission — failing to respond to a problem because the automation did not flag it. The system said nothing, so nothing was done.
- Errors of commission — following an automated directive despite contradictory information available from more reliable sources, because the operator either did not check or discounted what they found.
Raja Parasuraman and Victor Riley's Humans and Automation: Use, Misuse, Disuse, Abuse (Human Factors, 1997) gave the field its other durable frame: over-reliance is misuse, neglect of a useful aid is disuse, and both are failures. This is the origin of the inverted-U I use in the blueprint, and of John Lee and Katrina See's later Trust in Automation: Designing for Appropriate Reliance (2004), which is where the language of calibrated trust comes from. The target has never been maximal trust and has never been minimal trust. It is trust that tracks the actual reliability of the thing in front of you, which is a discrimination problem, and I will come back to that word because it determines how the examination has to be scored.
The medical evidence
Goddard, Roudsari and Wyatt, Automation bias: a systematic review of frequency, effect mediators, and mitigators (JAMIA, 2012), is the review that establishes the pattern in clinical decision support: overall performance usually improves, and the new errors the system introduces usually go unrecognised. Both halves of that sentence are true simultaneously, which is why arguments about clinical AI go round in circles — the advocate quotes the first half and the sceptic quotes the second, and neither is wrong.
Lyell and Coiera, Automation bias and verification complexity: a systematic review (JAMIA, 2017), screened 890 papers to 40 and found automation bias concentrated in single tasks with high verification complexity, typically diagnosis rather than monitoring — contradicting the received human-factors view that it is a multitasking phenomenon. Their conclusion is that automation bias tracks cognitive load, and that mitigation should therefore target load reduction.
Povyakalo, Alberdi, Strigini and Ayton, How to Discriminate between Computer-Aided and Computer-Hindered Decisions (Medical Decision Making, 2013), is the study I would put in front of any minister of health being sold a national screening deployment. Fifty professionals read 180 mammograms with and without computer-aided detection. The original analysis found no significant average effect. The reanalysis found that CAD raised sensitivity by about 0.016 for the 44 least discriminating readers on 45 relatively easy, mostly CAD-detected cancers — and lowered sensitivity by 0.145 for the 6 most discriminating readers on the 15 relatively difficult cancers.
Dratsch and colleagues, Automation Bias in Mammography (Radiology, 2023), gave 27 radiologists 50 mammograms with purported AI-generated BI-RADS categories, incorrect on 12 of the 40 in the test set. Incorrect suggestions impaired performance across the whole experience range, from inexperienced to very experienced. The inexperienced were hit hardest; the very experienced were hit too.
Tschandl and colleagues, Human–computer collaboration for skin cancer recognition (Nature Medicine, 2020), reached the two-sided version of the same conclusion: good AI support improved accuracy beyond either party alone, with the least experienced gaining most — and faulty AI misled the entire spectrum of clinicians, experts included.
Gaube and colleagues, Do as AI say: susceptibility in deployment of clinical decision-aids (npj Digital Medicine, 2021), ran a design I find quietly devastating. Physicians received chest radiographs with diagnostic advice — all of it written by human experts, some of it labelled as coming from an AI. Radiologists rated advice as lower quality when it carried the AI label; physicians with less task expertise did not. But diagnostic accuracy was significantly worse when the advice was inaccurate regardless of the label. The scepticism the experts expressed about the source did not translate into protection from the content.
Budzyń and colleagues, Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy (Lancet Gastroenterology & Hepatology, 2025), looked at four Polish centres in the ACCEPT trial and compared unassisted colonoscopy in the three months before AI was introduced with unassisted colonoscopy in the three months after. Adenoma detection rate fell from 28.4% (226 of 795) to 22.4% (145 of 648) — an absolute difference of −6.0 percentage points in procedures where no AI was present at all.
This is observational, it is a before-and-after within a trial context, and it has every confounder that design implies; the accompanying commentary says so. But it is the first substantial real-world signal that the underlying human skill degrades, and it degrades in the direction and on the timescale that would matter to us. Six points of adenoma detection is not a subtle finding.
Vaccaro, Almaatouq and Malone, When combinations of humans and AI are useful: a systematic review and meta-analysis (Nature Human Behaviour, 2024), pooled 106 experimental studies and 370 effect sizes. On average, human–AI combinations performed significantly worse than the better of human or AI alone. Losses were concentrated in decision-making tasks; gains were in content creation. Where humans outperformed the AI, the combination gained; where the AI outperformed humans, the combination lost.
And Krügel, Ostermaier and Uhl, ChatGPT's inconsistent moral advice influences users' judgment (Scientific Reports, 2023), established something with direct consequences for how we assess: users were influenced by inconsistent advice even when they knew it came from a chatbot, and they underestimated how much they had been influenced.
Part 3 — Why teaching literacy produces the problem it claims to solve
Here is the distinction the whole commitment turns on, stated as precisely as I can manage.
AI literacy is knowledge about the system. What a language model is. Why it produces fluent text without a model of truth. What a hallucination is and roughly how often it happens. How to write a better prompt. Which tasks it is good at. This is genuinely useful, it is teachable in a day, and it is what almost every clinical AI curriculum currently on offer consists of.
AI discernment is a behaviour in the presence of the system. It is the act of forming your own impression before you look; of noticing that the paragraph in front of you is internally consistent and externally wrong; of paying the verification cost when you are tired, behind, and the answer looks right. It is not knowledge. It is a performance, executed under load, repeatedly, for years.
The Qazi trial is the cleanest available demonstration that the first does not produce the second. Twenty hours of the first, and the second was absent when it was needed. And there are three mechanisms in the literature that explain why, none of which is fixed by more of the first.
Mechanism one: cognitive offloading
The availability of a plausible answer reduces the effort invested in generating your own. This is the mechanism the trial's authors invoke and it is the one clinicians recognise immediately when it is described to them, usually with a slightly guilty expression. It is not laziness in any morally interesting sense. It is the same adaptive process that lets you stop doing mental arithmetic when there is a calculator on the desk, and it is normally a good thing. It becomes dangerous only when the calculator is wrong sometimes and you have no cheap way of telling which times.
Mechanism two: the narrative surface
Qazi and colleagues make a point in their introduction that I think is underrated, and it distinguishes this generation of tools from everything the older literature studied. A conventional model outputs a discrete classification with a confidence score: malignant, 92%. That is a claim you can hold at arm's length. It announces itself as a machine output and it carries its own uncertainty on its face. A language model outputs a narrative — differentials with supporting and opposing evidence, a recommended next step, a caveat, a hedge in the right place. It has the surface features of clinical reasoning. It reads like a good registrar.
That surface is doing work. It is not merely that the output is persuasive; it is that the output mimics the very signals clinicians use to judge whether reasoning is sound. We are trained to trust an argument that considers alternatives and states its evidence. Here is a system that generates the form of that reliably and the substance of it only usually.
Mechanism three: fluency lowers friction
This is the mechanism I find most uncomfortable, because it implicates my own teaching. The trial found the significant decrement among physicians using language models at least weekly. The infrequent users' point estimate was nearly identical, so I will not claim the trial demonstrates that frequent use is worse — the honest reading is that the trial was underpowered to distinguish them. But it is consistent with a mechanism that is independently plausible: the more fluent the user, the lower the friction, the faster the consult, the less deliberate the reading, and the more the output slides into the reasoning unexamined.
The economics, stated plainly
Put the three mechanisms together with Lyell and Coiera's finding that automation bias tracks verification complexity, and the shape of the problem is economic rather than epistemic.
Consider what it costs a clinician to verify a plausible AI-generated differential in a district hospital at eleven at night. Two minutes at least, often ten. Attention taken from a queue. A working memory already carrying four other patients. And what does verification usually buy? Confirmation that the answer was fine. The expected return on any individual act of checking is low, because the tool is usually right — which is exactly the condition under which automation bias flourishes. A tool that was wrong half the time would train scepticism by itself, for free, in a week. A tool that is wrong three per cent of the time trains the opposite, and the three per cent is where the patients are.
Part 4 — The literature nobody designing an AI curriculum wants to cite
I have now made the case that scepticism must be trained. Before describing how, I have to deal with the evidence that training it does not work, because it exists, it is good, and I have not seen it cited once in a clinical AI curriculum document.
Taught debiasing does not transfer
The medical education community spent a decade on this in the context of ordinary diagnostic error — anchoring, availability, premature closure — long before AI. The intervention was cognitive forcing strategies: teaching clinicians to recognise the situations in which a given bias arises and to deliberately apply a counter-routine.
Sherbino, Kulasegaram, Howey and Norman, Ineffectiveness of cognitive forcing strategies to reduce biases in diagnostic reasoning: a controlled trial (CJEM, 2014), allocated 191 senior medical students on a four-week emergency medicine rotation to cognitive forcing strategy instruction or control, and tested them at the end of the rotation. There was no difference in the rate of diagnostic error between the groups. An earlier exploratory study by the same group had pointed the same way.
That result is not isolated. The pattern across the debiasing literature in medicine is that instruction in bias recognition reliably improves people's ability to name biases and reliably fails to improve their diagnostic accuracy. Knowing the name of the trap does not keep you out of it.
And from the aviation side, Skitka, Mosier, Burdick and Rosenblatt, Automation Bias and Errors: Are Crews Better Than Individuals? (International Journal of Aviation Psychology, 2000), tested training that focused explicitly on automation bias and its associated errors. It reduced commission errors and did not reduce omission errors.
What this means for commitment two
Taken at face value, this literature says: writing "scepticism is trained explicitly" into a founding document is close to writing a wish. If I stopped here, the intellectually honest move would be to strike the commitment.
I do not think that is the right conclusion, and here is the argument for why — which is also the argument for what the commitment has to be changed into.
The interventions that failed were instructional. A course. A strategy. A named routine to be recalled and applied under load, by an individual, unprompted, in an environment that neither required it nor checked for it. The interventions that have worked in adjacent domains were structural and procedural: they changed the task, or they changed what the person was accountable for, or they drilled the behaviour to automaticity against a standard.
Part 5 — What the evidence says does work
Three things have support. None is a lecture.
One: cognitive forcing as structure, not as strategy
The distinction between Sherbino's failed intervention and this one is subtle and it is everything. Sherbino taught people a strategy to apply. The alternative is to build the forcing function into the task so that no recall, no discipline and no virtue is required.
Buçinca, Malaya and Gajos, To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making (Proceedings of the ACM on Human–Computer Interaction, 2021), tested interventions that structurally interrupt the path of least resistance — most relevantly, requiring the person to commit to their own answer before the AI's recommendation is revealed. Their cognitive forcing interventions significantly reduced overreliance, where simply adding explanations to the AI's recommendation did not.
Their theoretical framing is worth stating because it explains why explanation-based approaches keep disappointing: people rarely engage analytically with each individual recommendation and explanation. They develop a general heuristic about whether to follow this system, and then apply the heuristic. A better explanation feeds the heuristic. Only a structural interruption reaches the analytical process at all.
Two: simulation-based mastery learning
If discernment is a performance rather than a knowledge state, it belongs in the tradition of procedural skill acquisition, where the evidence is far more encouraging than in the debiasing tradition.
McGaghie, Issenberg, Cohen, Barsuk and Wayne's meta-analytic comparison of simulation-based medical education with deliberate practice against traditional clinical education found a pooled effect size of 0.71 (95% CI 0.65–0.76) in favour of simulation with deliberate practice, across the 14 of 3,742 screened studies that met inclusion criteria, and their critical review of simulation-based mastery learning with translational outcomes (Medical Education, 2014) documents transfer to patient-level results — central line infections, for the canonical example.
The operative words are deliberate practice and mastery. Not exposure to a simulator. Repeated attempts against a pre-set standard, with immediate specific feedback, continuing until the standard is reached, with the time to reach it allowed to vary between learners rather than the standard varying. That is a different and more expensive thing than what most institutions mean when they say simulation, and it is the only version with this evidence behind it.
Three: accountability
Skitka, Mosier and Burdick's accountability study found that making participants accountable — for their overall performance or specifically for their decision accuracy — lowered rates of automation bias. This is the cheapest of the three interventions and the one most easily eroded, because accountability that is announced but never enacted decays within about one rotation.
Institutionally it means the countersignature is not paperwork, the audit returns to a named person, and somebody senior occasionally asks a clinician why they accepted a particular recommendation — not punitively, and not rarely enough to be forgettable.
Part 6 — The design
This is the part I would want criticised on specifics. Everything above is argument; what follows is a set of decisions, several of which I hold loosely and one or two of which I suspect are wrong in ways I cannot yet see.
6.1 The independent-impression rule
The rule. Before any AI output is visible, the clinician records their own impression: a leading diagnosis, two alternatives, and the single finding that would most change their mind. Only then does the output unlock.
This is Buçinca's cognitive forcing function, implemented as a workflow constraint rather than as advice. In simulation it is enforced by the platform — the output is behind a control that does not respond until the impression field is complete. In the workplace it is enforced by protocol and by the record, which is weaker, and I will come back to that.
Three failure modes I would design against from the start.
The hedged impression. The obvious defeat is to write something so vague it cannot be contradicted — "query sepsis, query cardiac, query other" — preserving the option to agree with whatever appears. This converts the rule into a keystroke. So the impression is scored for specificity and commitment, and an impression that does not commit is scored as absent, not as partial. Learners are told this explicitly in the first hour, because the point is not to catch them out.
The retrospective edit. If the interface permits revision of the impression after the output is revealed, the data is worthless and so is the drill. The impression is timestamped and locked. This is a technical requirement, not a policy one.
Ritualisation. The most likely long-run failure is that the rule survives as a form and dies as a habit — a box filled in at speed by someone who has already decided to read the output first. I do not have a clean defence against this. The partial defences are that the impression is periodically scored rather than merely collected, and that the workplace-based assessment described in the Kirkpatrick companion looks at the content of impressions and not their presence.
6.2 The seeded-error bank
Every simulation encounter runs on a sandbox where we control what the model appears to say. A proportion of encounters contain a seeded clinically significant error.
The taxonomy. This is the part that has to be Kenyan and specific, per commitment five, and it has to be weighted towards omission, per Skitka. Working categories:
| Category | Example | Type |
|---|---|---|
| Dose or interval wrong for organ function | A standard dose in a patient with an eGFR of 22 | Commission |
| Omitted local differential | A febrile returning-traveller differential with no malaria, or a lymphadenopathy differential with no TB | Omission |
| Omitted red flag | A back pain assessment that never asks about bladder function | Omission |
| Fabricated or misattributed citation | A confident reference to a guideline that does not say that | Commission |
| Guideline correct elsewhere, wrong here | A recommendation presupposing an investigation unavailable at a Level 4 facility | Commission |
| Nomenclature collision | A drug name that maps to a different agent in our supply chain | Commission |
| Confidently normal reading of an abnormal value | A potassium of 6.1 characterised as "mildly raised, monitor" | Commission |
| Silent scope error | A differential that answers the question asked while ignoring the more important question | Omission |
| Skin-tone-dependent misread | A cellulitis or pressure-area assessment degraded on darker skin | Commission |
| Language-transfer error | A history taken in Kiswahili or Dholuo, rendered into English, with the compounding error preserved | Commission |
The base rate problem, which I do not think has a clean answer.
In the Qazi trial the seeded-error rate was 50% of cases in the treatment arm. In the blueprint I said roughly one in three, varied so learners cannot game it. Both of those are wildly higher than the real-world rate of clinically significant error in a competent contemporary model, which — depending on task, model and how you define significant — is plausibly in the low single figures per cent.
That gap is a genuine design tension. Train at a high rate and you build a vigilant reflex, but you also train a criterion calibrated to a world that does not exist, and you risk producing under-trust: the clinician who rejects the useful retinopathy screen and misses the retinopathy. Train at the true rate and most learners will complete an entire simulation block without encountering a single error, which trains nothing at all and wastes very expensive faculty time.
My resolution, offered as a decision rather than a solution:
- The training rate is deliberately high and deliberately unstable — varied between 20% and 40% across blocks, never announced, never fixed. High enough to generate practice repetitions; unstable enough that no base rate can be learned and applied as a heuristic, which is precisely the failure Buçinca describes.
- The assessment rate is set separately, disclosed as a design choice in the assessment blueprint, and held constant within a cohort for fairness. Candidates are told that some proportion of stations contain errors and are not told which or how many.
- The recertification rate is set from our own audit data — whatever our incident reporting and case review say the real local rate is — because by then the objective is calibration rather than acquisition.
- Under-trust is measured as an outcome, not assumed away. See the scoring below, which is the only part of this design that takes it seriously.
6.3 Scoring: why "error-catch rate" is the wrong metric
In the blueprint I wrote that error-catch rate would be a primary assessed outcome, pass or fail. Having thought about it properly, that formulation is wrong, and wrong in a way that would have produced a perverse cohort.
Error-catch rate alone is trivially gameable. A candidate who challenges every AI output catches 100% of seeded errors. They also reject every correct recommendation, and if the examination scores only catches, they pass at the top of the cohort. We would have certified, and then released onto wards, the clinician who has learned that the safe answer in an examination is always to disagree with the machine — which is Parasuraman's disuse, and which the blueprint's own inverted-U identifies as a real failure mode with real patients behind it.
The correct frame is the one the aviation and psychophysics literature has used for decades: this is a signal detection problem, and it has two independent parameters.
- Sensitivity (discrimination) — how well the candidate separates erroneous output from sound output. This is the skill.
- Criterion (bias) — how readily they say "error" regardless. This is a disposition, and it can be moved without any change in skill at all.
A candidate is characterised by both. High sensitivity with a sensible criterion is discernment. High catch rate achieved by a permissive criterion is not.
What I would therefore score, per station:
- Detection — was the seeded error identified? (Hit)
- False alarm — was sound output challenged as erroneous? (False alarm, scored with equal seriousness)
- Action — was the right thing then done? Catching an error and proceeding anyway is a failure, and it is a distinct failure from not catching it.
- Reasoning — is the stated basis for the challenge correct? A candidate who rejects a good recommendation for a bad reason and happens to be right has demonstrated nothing.
- Impression integrity — was an independent impression recorded, specific, and committed, before the output was revealed?
6.4 The interprofessional station
Commitment eight says interprofessional wherever the work is interprofessional, and the space between a nurse and a consultant is where these failures actually live.
So one station in every assessment cycle is built like this. A confederate playing a senior clinician has already accepted an AI-generated recommendation that contains a seeded error. The candidate — a nurse, a clinical officer, a pharmacist, a junior doctor — has the information needed to detect it. The assessed behaviour is not detection. It is speaking up.
The vehicle is SBAR, which exists for exactly this: a structure that lets a junior person deliver an unwelcome assessment to a senior one without having to be brave in the moment. The recommendation clause is where the challenge goes, and having a slot to put it in is most of the battle.
Two design commitments follow, and both cost money:
- The consultant is assessed too, on the receiving end. A cohort of nurses trained to challenge, released into a hospital of consultants who have never practised being challenged, is an experiment in professional attrition. Commitment nine — faculty are certified and their teaching is observed — is not separable from this. You cannot examine a reflex you do not have.
- The mixed-cadre group is not a scheduling convenience. The Level 1 common core is taught in mixed groups deliberately, and this station is why.
6.5 Adversarial rounds
A recurring seminar in which the learner's job is to break the model on a case from their own practice and explain the failure mechanism to the group. Three things this does that simulation does not:
- It builds a local catalogue of failure modes — the ones that occur in Kenyan practice, on Kenyan patients, with the drugs and investigations we actually have. Nobody else will build this for us.
- It shifts the learner from the receiving end to the adversarial end, which is where the useful intuitions live.
- It gives us a continuously refreshed source of new seeded errors from real practice, which is the only defence I have against the objection in Part 8 that our seeded errors are limited to the failures we can imagine.
The output is a written failure-mode entry, countersigned, which is also the learner's portfolio product under commitment seven.
6.6 Where this sits in the programme
| Level | Discernment content | Assessment |
|---|---|---|
| Level 1 common core (12 h, all cadres) | The independent-impression rule; the modalities; one seeded-error simulation | Formative; impression integrity recorded |
| Level 2 | Full seeded-error simulation block; error taxonomy by cadre | Summative AI-OSCE, conjunctive rule |
| Level 3 | Interprofessional challenge station; adversarial rounds begin | Summative, plus countersigned failure-mode entry |
| Level 4 | Supervising others' AI use; assessing a junior's discernment | Workplace-based, multi-source |
| Fellowship | Building the local error bank; running adversarial rounds | Teaching observed and certified |
| Recertification (2 yr) | Re-tested at audit-derived base rate | Summative; unassisted performance also measured |
The one row I would defend hardest is the last. Recertification is where a reflex that has quietly decayed becomes visible, and it is the row most likely to be cut for cost.
Part 7 — Measuring whether it held
The full apparatus is in the Kirkpatrick companion, so this is only what is specific to discernment.
The unassisted-performance audit. Budzyń and colleagues did something in colonoscopy that we should copy deliberately rather than discover accidentally: they measured the clinician without the tool, before and after exposure to it. That is the only design that detects deskilling, and it is not expensive if it is planned. So: a proportion of workplace observations are conducted on cases where AI is not used, and the cohort's unassisted diagnostic performance is tracked over time as a balancing measure.
Decay. The independent-impression rule is my candidate for the fastest-decaying element, because it is a small effortful behaviour with no immediate reward, executed in private, under time pressure. I would measure impression integrity at three and twelve months, expect it to fall, and treat the shape of the fall as the thing that determines where the booster goes. If it does not fall, I was wrong and should say so.
The behavioural indicator I care most about. Not the error-catch rate at examination. The rate at which clinicians, in real practice, record an impression that differs from the AI's leading suggestion. A cohort in which that rate approaches zero has either achieved perfect agreement with a perfect tool or stopped thinking, and the two are distinguishable only by the case review that follows.
Part 8 — What this does not do
Every security document should have this section and so should every curriculum.
It rests on a small evidence base at the point where it is most specific. The direct evidence that AI-literacy training fails to prevent automation bias in physicians is, at present, one randomised trial of 44 people in one city with one model. The rest of my case is convergent evidence from adjacent modalities and from aviation, which is strong for the general phenomenon and silent on the particulars of the intervention I am proposing.
Transfer is assumed, not demonstrated. Simulation-based mastery learning transfers for procedural skills — central lines, lumbar punctures, resuscitation. Discernment is not a procedure; it is closer to a habit of mind, and habits of mind are exactly the category where the debiasing literature says transfer fails. My argument is that treating it as a procedure, with a structural forcing function, is what moves it into the category where transfer works. That argument is plausible and untested. It is the central bet of the design and it may be wrong.
Our seeded errors are the errors we can imagine. This is the objection I find hardest. A bank built by clinicians contains the failures clinicians anticipate; the dangerous model failures are, almost by definition, the ones nobody thought to seed. Adversarial rounds are a partial answer because they harvest real failures from practice, but they will always lag. There is no version of this design that trains against an unknown failure mode, and any claim otherwise would be marketing.
The curriculum has a half-life. Specific failure modes are properties of specific model versions. The taxonomy in 6.2 is a snapshot of 2026 and a substantial part of it will be obsolete within two years, either fixed or replaced by something stranger. Only the procedure is durable — impression first, verify, document, act — and even that is a bet on the shape of the tools rather than a law.
Signal-detection scoring assumes a stable criterion, and we assess rested candidates. Criterion shifts under fatigue, and fatigue is the condition under which the failure actually occurs. Examining people at ten in the morning in a simulation centre measures the capacity, not the behaviour at three a.m. The workplace-based component exists to close that gap and closes it only partially.
It may be that interface design does more than education can. The Qazi authors say as much — provenance cues, uncertainty display, bias-aware interfaces. If a well-designed system produces more discernment than a well-designed curriculum, then a serious institution's job is partly to write procurement standards rather than to teach, and I would want to know that rather than defend my own turf. Commitment two is an educational claim, and educational claims about problems with engineering solutions have a poor history.
And the whole thing might be unpopular enough to be abandoned. Buçinca's trade-off is not a footnote. The design deliberately makes work slower and more irritating, and the people subjected to it will rate it accordingly. Institutions delete what scores badly.
What would refute this
Stated in advance, because a commitment that cannot fail is not a commitment.
- The direct test. Cohorts trained under this design should show a higher error-catch sensitivity at twelve months than cohorts given a conventional twenty-hour AI-literacy course, with no worse a false-alarm rate. If they do not, the design is wrong.
- The transfer test. Impression integrity in workplace observation at twelve months should exceed what an untrained comparator does spontaneously. If the simulation performance is excellent and the ward behaviour is identical to control, I have built an examination rather than a competence.
- The deskilling test. Unassisted diagnostic performance in trained cohorts should not decline relative to baseline. If it declines the way Budzyń's endoscopists' did, the training has not protected the underlying skill and something more than curriculum is required.
- The decay test. If impression integrity is stable at twelve months without boosters, my central hypothesis about decay is wrong, and the recertification interval and the booster schedule should both be relaxed. That would be good news and I would publish it as readily as the bad.
All four are answerable with the evaluation design already specified, and all four should be registered before the first cohort is taught.
What I would write into the founding documents
Compressed, so it fits on one page of an operations manual.
- Commitment two is amended to read: scepticism is trained explicitly, built into the workflow structurally, and assessed behaviourally. The middle clause is not optional and its absence is what the evidence predicts would sink it.
- The independent-impression rule is a structural constraint, not an instruction. The AI output does not unlock until a specific, committed impression is recorded and timestamped. Impressions cannot be edited after the output is revealed.
- A hedged impression is scored as absent. Stated to learners in the first hour.
- No less than 40% of seeded errors are errors of omission. Because training already handles the other kind.
- The training error rate is unstable by design and never announced; the assessment rate is fixed within a cohort, disclosed in the assessment blueprint, and set by the standard-setting panel; the recertification rate is derived from our own audit data.
- No candidate is ever assessed on catch rate alone. Sensitivity, false-alarm rate, action and reasoning are scored separately, the rule is conjunctive, and compensation between them is not permitted. Under-trust is a failure mode with a pass mark attached to it.
- Faculty are assessed on the receiving end of challenge before they are certified to teach the challenging.
- Unassisted performance is measured as a balancing outcome and published whichever way it moves.
- Satisfaction data is never permitted as evidence that this element works, and a fall in satisfaction scores after the forcing function is introduced is pre-declared as an expected finding rather than an adverse one.
Coda
The reason this commitment is nine words long and the post is not is that the nine words conceal a decision most curricula never make.
You can teach a clinician what a language model is in an afternoon, and they will leave knowing more and behaving identically. The Qazi trial is a photograph of that afternoon's consequences: forty-four physicians who could all have defined a hallucination, sitting in front of one, and writing it down.
What has to be built instead is uncomfortable in a specific way. It slows the clinician down. It makes them commit before they are ready. It marks them on the recommendations they wrongly rejected as well as the ones they wrongly accepted. It requires a nurse to contradict a consultant in a room where both are being watched, and it requires the consultant to be examined on how they take it. None of that is popular, and the evidence says plainly that it will be rated worse than the lecture it replaces.
The reflex in the title is resented because it costs something every single time and pays off almost never — until the night it is the only thing standing between a fluent, confident, well-formatted paragraph and a patient. Skills like that do not survive on enthusiasm. They survive because an institution wrote them down, drilled them, examined them, and was willing to fail people who did not have them.
That is what the second clause of commitment two is for. Everything else in this post is an argument about how to make it enforceable.
If you want the rest of the design: the full blueprint sets out the institution, the five tracks and the five gated levels; Borrowed From an Art School traces where the competency framework came from; One Hidden Error covers the OSCE and the AI-OSCE that this post's stations sit inside; The Angoff Panel covers where the pass marks in Part 6.3 would come from; Measuring What Actually Matters covers the evaluation apparatus behind Part 7; Four Letters From a Submarine covers SBAR, which is the vehicle for the interprofessional station; and The Law Is Part of the Architecture covers the Kenyan legal frame around the simulation data. Other writing is in the archive, and things I have built are under demos and lessons.
References
The trial the argument rests on
- Qazi, I. A. et al. (2025). Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial. NEJM AI. Preprint: medRxiv 2025.08.23.25334280. Registered as NCT06963957.
Automation bias — the foundational work
- Skitka, L. J., Mosier, K. L. and Burdick, M. (1999). Does automation bias decision-making? International Journal of Human–Computer Studies 51(5):991–1006. The omission/commission distinction.
- Skitka, L. J., Mosier, K. and Burdick, M. D. (2000). Accountability and automation bias. International Journal of Human–Computer Studies 52(4):701–17.
- Skitka, L. J., Mosier, K. L., Burdick, M. and Rosenblatt, B. (2000). Automation Bias and Errors: Are Crews Better Than Individuals? International Journal of Aviation Psychology 10(1):85–97. Training reduced commission but not omission errors.
- Parasuraman, R. and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors 39(2):230–53.
- Lee, J. D. and See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors 46(1):50–80. The origin of calibrated trust.
Automation bias in clinical settings
- Goddard, K., Roudsari, A. and Wyatt, J. C. (2012). Automation bias: a systematic review of frequency, effect mediators, and mitigators. JAMIA 19(1):121–7.
- Lyell, D. and Coiera, E. (2017). Automation bias and verification complexity: a systematic review. JAMIA 24(2):423–31. Automation bias tracks cognitive load and verification complexity.
- Povyakalo, A. A., Alberdi, E., Strigini, L. and Ayton, P. (2013). How to Discriminate between Computer-Aided and Computer-Hindered Decisions: A Case Study in Mammography. Medical Decision Making 33(1):98–107. The null average concealing help to weak readers and harm to strong ones.
- Dratsch, T. et al. (2023). Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology 307(4):e222176.
- Tschandl, P. et al. (2020). Human–computer collaboration for skin cancer recognition. Nature Medicine 26:1229–34. Faulty AI misleads the entire spectrum, experts included.
- Gaube, S., Suresh, H., Raue, M. et al. (2021). Do as AI say: susceptibility in deployment of clinical decision-aids. npj Digital Medicine 4:31.
- Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterology & Hepatology 10(10):896–903. ADR 28.4% → 22.4% in unassisted procedures.
- Vaccaro, M., Almaatouq, A. and Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour 8:2293–2303.
- Krügel, S., Ostermaier, A. and Uhl, M. (2023). ChatGPT's inconsistent moral advice influences users' judgment. Scientific Reports 13:4569. Users underestimate how much they are influenced.
Why instruction alone fails, and what works instead
- Sherbino, J., Kulasegaram, K., Howey, E. and Norman, G. (2014). Ineffectiveness of cognitive forcing strategies to reduce biases in diagnostic reasoning: a controlled trial. CJEM 16(1):34–40. And the earlier exploratory study (2011).
- Buçinca, Z., Malaya, M. B. and Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proceedings of the ACM on Human–Computer Interaction 5(CSCW1), Article 188. Preprint. Includes the performance/preference trade-off.
- McGaghie, W. C., Issenberg, S. B., Cohen, E. R., Barsuk, J. H. and Wayne, D. B. (2011). Does Simulation-Based Medical Education with Deliberate Practice Yield Better Results than Traditional Clinical Education? Academic Medicine 86(6):706–11. Pooled effect size 0.71 (95% CI 0.65–0.76).
- McGaghie, W. C., Issenberg, S. B., Barsuk, J. H. and Wayne, D. B. (2014). A critical review of simulation-based mastery learning with translational outcomes. Medical Education 48(4):375–85.
Kenyan and framework context
- OpenAI and Penda Health, Pioneering an AI clinical copilot; the underlying real-world study; and the critical reading in STAT News.
- Dakan, R. and Feller, J., AI Fluency Framework and the Practical Summary Document. The Clinical 4Ds are an adaptation; the licence terms are set out in Borrowed From an Art School.
- WHO, Ethics and governance of AI for health: guidance on large multi-modal models; AAMC, Artificial Intelligence Competencies Across the Learning Continuum.