I have spent my working life at an unusual junction. I trained in medicine and I have operated. I have also built systems — engineering, information technology, and latterly artificial intelligence. For most of my career those two halves sat in separate rooms of my head, and I was periodically irritated by how little the people in one room knew about what the people in the other were doing.
That separation is no longer sustainable, and it is no longer merely irritating. It is becoming dangerous.
Artificial intelligence has arrived in Kenyan clinical practice ahead of any coherent plan for teaching clinicians how to use it. It arrived, as these things do, informally: a clinical officer in a busy outpatient department typing a differential into a chatbot on a personal phone between patients; a registrar drafting a discharge summary with a model that has never seen a Kenyan formulary; a consultant checking a dose at two in the morning against a system nobody has validated against our epidemiology.
None of these people were taught to do this. None of them were taught not to. Nobody assessed whether they could tell a good answer from a confident wrong one.
This post sets out what I would build to close that gap, and — more importantly — how I would prove that it worked.
Download the full blueprint
The complete document is forty-four pages: the case, the institutional architecture, the pedagogy, the curriculum across five professional tracks and five gated levels, the quality system, the team specified post by post, a thirty-six month sequence, a risk register, and seventy-one references. This post is the argument. The PDF is the plan.
⬇ Download PDF (376 KB, 44 pages, 8 figures)
The position I am arguing from
I want to be clear at the outset, because it determines everything that follows.
I do not think artificial intelligence is going to practise medicine. I think it is going to be a tool that people who practise medicine will carry. The right analogy is not the arrival of a new colleague. It is the arrival of the stethoscope, the ultrasound probe, the pulse oximeter, the laparoscope. Each of those was, at introduction, greeted by a mixture of exaggerated hope and exaggerated fear. Each turned out to be a tool: enormously useful in trained hands, useless or harmful in untrained ones, and never a substitute for the judgement of the person holding it.
A clinician in 2026 walks into a consulting room carrying a quiver. In it are history-taking, examination, pattern recognition earned over thousands of patients, the laboratory, imaging, the formulary, guidelines, colleagues, and the accumulated literature. AI is another arrow. It is a peculiar arrow — it is fluent, it is fast, it is confidently wrong in ways that no previous tool has been, and it flatters the user — but it is an arrow, not an archer.
Kenya is already generating the evidence
Kenya is not a bystander in clinical AI. It is one of the places where the most consequential real-world evidence is being produced.
In 2025, a study across fifteen primary-care clinics in Nairobi examined nearly 40,000 patient visits in which clinicians had access to an AI decision-support tool integrated into the electronic record. It reported a 16% relative reduction in diagnostic errors and a 13% relative reduction in treatment errors among clinicians who engaged with the tool. That is a substantial signal, and it was generated here, on our case mix, by our clinicians.
But read the same body of work carefully and a second finding sits alongside the first, less comfortable and far more instructive: the benefit was concentrated among clinicians who used the tool well. Uptake and quality of engagement varied enormously between individuals. The technology did not distribute its benefit evenly. It distributed it to the people who knew how to use it.
That is the case for this institution in a single sentence.
The finding that should change how we teach this
Here is where it gets uncomfortable, and where most AI-in-medicine curricula currently being written are going to fail.
A randomised trial published in 2025 took physicians who had already completed twenty hours of AI-literacy training — covering model capabilities, prompt engineering, and critical evaluation of AI output — gave them clinical cases, and exposed half of them to deliberately erroneous output from a large language model.
They deferred to it.
Prior literacy training did not protect them. Consultation was voluntary; they retained full autonomy to accept, modify or reject. They accepted. The mechanism is well described in the cognitive literature: cognitive offloading, in which the availability of a plausible answer reduces the effort the clinician invests in generating their own.
I have thought about that result more than any other in this field, and it is the axis on which I would design the entire curriculum.
The target is not maximal trust, and it is not minimal trust.
Under-trust is not a safe default. The clinician who refuses the screening tool and misses the retinopathy has harmed a patient just as surely as the one who accepted a wrong answer. But over-trust is the failure mode the evidence says we are heading for, and it is the one that fluency actively produces.
How I would train against it:
The independent-impression rule. Drilled from the first hour and enforced in every simulation: form and record your own clinical impression before you look at the AI output. Not after. This single behavioural constraint does more to prevent anchoring than any amount of exhortation, because it makes the clinician's own reasoning a fixed point rather than something that gets quietly revised.
Seeded-error simulation. Every simulation encounter runs on a sandbox where we control the AI's output. Roughly one in three contains a seeded clinical error — and the rate is varied so learners cannot game it. The errors are not cartoonish. They are the errors these systems actually make: a plausible dose that is wrong for renal impairment; a differential that omits the tropical diagnosis; a confident citation to a paper that does not exist; a guideline recommendation correct for a European population and wrong for ours; a drug name that is right internationally and refers to something else locally.
Error-catch rate as a primary assessed outcome. Not a formative nicety. Pass or fail. A candidate who does not catch seeded errors at the standard-set threshold does not progress, however fluent their prompting.
Adversarial rounds. A recurring format in which the learner's explicit job is to break the AI on a case from their own practice, then explain the failure mechanism to the group. This builds a shared institutional catalogue of failure modes specific to our context, and inoculates against the deference that fluency induces.
Skill-decay monitoring. Recertification at two years, and not a formality. If a cohort's unassisted diagnostic performance is deteriorating, that is a finding about our training — and one we would publish.
Why this would be the first institution of its kind anywhere
I want to state this plainly and then defend it, because it is a strong claim.
What I am describing would be a one-of-a-kind institution the world over. It does not currently exist anywhere on earth.
There are, of course, excellent things adjacent to it. Mount Sinai runs a mid-career data and AI skills programme for its faculty. Harvard Medical School and the Harvard Chan School run executive and continuing-education courses on AI in clinical medicine and its implementation. The University of Florida has built an AI-in-medicine curriculum. The AAMC is developing AI competencies across the medical education continuum, and the proportion of North American medical schools incorporating AI into their curricula rose from 53% to 77% in a single year. Anthropic, working with Professors Rick Dakan and Joseph Feller, has released a genuinely excellent open framework and course series on AI fluency. Radiology and informatics fellowships exist.
Every one of those is a course, a fellowship, or a curricular component inside an existing school. Each serves one institution, one specialty, or one professional stratum. Each is, essentially, an elective.
What does not exist anywhere is this:
A permanent, national, publicly-accountable institution whose sole mandate is the clinical AI competence of an entire country's health workforce — every cadre, from the consultant surgeon to the ward nurse to the hospital administrator — delivering competency-gated certification tied to professional licensure, assessed by simulation and by workplace observation rather than by attendance, and publishing its own outcome data whether or not the outcomes are flattering.
Nobody has built that. Not the United States, not the United Kingdom, not Singapore, not Germany. The wealthy systems have distributed the problem across a thousand medical schools, professional colleges and vendor training programmes, each optimising locally, none accountable for the whole. That fragmentation is a function of their size and their institutional inertia, and they will be years unwinding it.
Kenya's position is different, and the difference is an advantage:
- Our regulatory councils are national and singular.
- Our CPD architecture is already mandatory and centrally administered — the Kenya Medical Practitioners and Dentists Council requires 50 CPD points per calendar year for retention on the register; the Nursing Council of Kenya operates its own points requirement for licence renewal.
- Our health system has a clean six-level structure from community units to national referral hospitals.
- We have a national AI strategy (2025–2030) that names health as a priority sector.
- We have a digital health statute — the Digital Health Act 2023 — establishing a Digital Health Agency and setting data principles, alongside the Data Protection Act 2019.
- And we have the rare and precious circumstance of a workforce that is adopting a technology right now, in real time, before habits have calcified.
I do not think that is grandiosity. I think it is an accurate reading of a narrow window. The institution I am describing would place Kenya at the spearhead of this development for the whole of humanity — not because we are wealthier or better resourced than others, but because we are more agile, because we are already generating the evidence, and because we would be willing to publish what we found. The countries that will follow us are the ones currently writing committee reports about it.
What I am not proposing
Three disclaimers, because the failure modes here are well known.
This is not a programme to make clinicians into machine-learning engineers. A very small number of our fellows will go deep into model architecture. The overwhelming majority of the workforce needs something quite different: the judgement to use a tool safely. Confusing those two is the commonest error in this field and it produces curricula that are simultaneously too technical to be useful and too shallow to be rigorous.
This is not a vehicle for any vendor. The institution must be able to teach against, criticise, and if necessary publicly fail any commercial system — including systems deployed in the hospitals it serves. That independence has to be structural, not aspirational.
This is not a substitute for clinical training. If a clinician's underlying medicine is weak, AI will not fix it; it will amplify the weakness and make it fluent. We are adding an arrow to a quiver. We are not issuing the quiver.
The institution
I would call it the Kenya Institute for Clinical Artificial Intelligence. Its mandate would be one sentence, and I would have it carved somewhere visible:
Every clinician in Kenya can use artificial intelligence safely, sceptically, and to the patient's benefit — and can prove it.
Note the last four words. They are the ones that make this an institution rather than a lecture series.
It would be established as a semi-autonomous training and standards body hosted by a major teaching hospital, with its own governing board, its own budget line, and — this matters more than anything else — the legal capacity to withhold a certificate. An institution that cannot fail a candidate is not a training institution. It is a conference.
The five functions are teaching; assessment and certification; simulation and sandbox; evaluation and research; and standards and advisory. They are shown in the figure at the top of this post.
Governance, and why it must be uncomfortable
Three specific irritants, built in deliberately, because an institution that is comfortable is an institution that has stopped checking itself:
The Office of Quality and Evaluation reports to the board, not to me. The people who measure whether the training works must not be paid by the people who deliver it, and must not be promotable by them. This is the single most important structural decision in the whole design, and the one most likely to be quietly eroded within three years if it is not written into the founding instruments.
Outcomes are published annually, including the bad ones. Pass rates, failure rates, dropout, modules that did not work, assessments that turned out to have poor discrimination, and any incident in which a trained clinician's use of AI contributed to patient harm. A training institution that only publishes its successes is generating marketing, not evidence.
A patient and public panel reviews the curriculum. Twelve lay members, properly remunerated, who read what we propose to teach clinicians about how to use AI on them. If we cannot explain a module to them, we do not understand it well enough to teach it.
The pedagogy I would insist on
This is the part I care about most and would be least willing to compromise on. Ten commitments, which I would want written into the founding documents so that a future director has to argue publicly to abandon them:
- We teach judgement, not tools. Test every module against one question: if the vendor disappeared overnight, would this teaching still be worth anything? If no, it is training, not education, and it belongs in a vendor manual.
- Scepticism is trained explicitly, and it is assessed. Not a lecture on limitations — a drilled reflex, like recognising a deteriorating patient.
- Nothing is taught that is not assessed, and nothing is assessed that was not taught.
- Attendance certificates are abolished. No award for having been present. This will make us unpopular and it is not negotiable.
- All teaching is case-based and Kenyan. Not a single vignette involving insurance codes, drugs we cannot obtain, or investigations we do not have.
- Simulation before patients, always. Uncontroversial for central lines. It should be uncontroversial here.
- The learner produces something. A logbook, a critique, an evaluation, a taught session — read and countersigned by a named senior person.
- Interprofessional wherever the work is interprofessional. Ward AI use is not a doctor problem or a nurse problem; the failure modes live in the handover between them.
- Faculty are certified, and their teaching is observed.
- We measure at Kirkpatrick 3 and 4, or we admit we do not know. Satisfaction scores are close to worthless.
The intellectual spine: the Clinical 4Ds
I would not invent a competency framework from scratch. There is a good one, it is well constructed, it is open, and reinventing it would be vanity.
The AI Fluency Framework — and its four core competencies, Delegation, Description, Discernment and Diligence — was developed by Professor Rick Dakan of Ringling College of Art and Design and Professor Joseph Feller of Cork University Business School, University College Cork, and elaborated into a course series in partnership with Anthropic, with support from Ireland's Higher Education Authority.
Before the competencies, the framework sets out three modalities of human–AI interaction. I find them unusually useful clinically, because they carry different risk profiles:
| Modality | Framework definition | Clinical instance | Dominant risk |
|---|---|---|---|
| Automation | AI performs a task independently on direct human instruction | Drafting a discharge summary from structured notes | Unreviewed output entering the record |
| Augmentation | AI and human co-define and co-execute a task iteratively | Working through a difficult differential | Anchoring; the clinician's own reasoning quietly revised |
| Agency | Human configures AI to perform future tasks independently, including for others | A triage assistant configured once and left running on a queue | Harm at scale, with no clinician in the room when it occurs |
Teaching clinicians to name which modality they are in — before they act — is one of the highest-yield twenty minutes in the whole common core. Almost all the serious failure modes I can construct involve someone operating in agency mode while believing they are in augmentation mode.
The framework's teaching course adds one further structural idea that maps onto clinical work almost too neatly. The four competencies are not a list; they are two loops. Delegation and Diligence form the outer loop — what you decide to hand over, and what you take responsibility for afterwards. Description and Discernment form the inner loop — how you ask, and how you judge what comes back — and it turns round many times within a single encounter. A clinician will recognise the shape immediately: consent-and-audit on the outside, history-and-examination on the inside.
Delegation, clinically, is the construction of a non-delegable list — the clinical acts that never leave your hands regardless of how good the tool becomes. My own list, offered as a starting point for argument rather than doctrine: obtaining consent; breaking bad news; the decision to operate; the final diagnosis committed to the record; the prescription; the signature. Around that hard core sits a much larger and genuinely negotiable territory. Teaching the boundary is the work.
Description, clinically, is something a clinician already knows how to do and does not know they know. A good prompt is a good handover. The SBAR structure a nurse uses to hand over a deteriorating patient is a better prompt template than anything in the prompt-engineering literature. What must be added is the context a Kenyan clinician takes for granted and a model does not: which level of facility, which formulary, which tests exist, what the local prevalence actually is. And one absolute rule, taught on day one and assessed: identifiable patient data does not go into a system you do not control.
Diligence, clinically, is documentation of what was AI-assisted and how; disclosure to patients and colleagues; data protection and consent within Kenyan law; incident reporting where AI contributed to harm; and the professional obligation to guard your own skills against decay. The point I would drive hardest: the signature is yours, the liability is yours, and no tool has ever taken responsibility for anything.
Bias, taught concretely or not at all
Generic teaching about algorithmic bias is nearly useless because it stays abstract. Taught locally, it bites:
- Dermatological and wound-assessment models trained predominantly on lighter skin, and what that does to a diagnosis of cellulitis or a pressure sore on a Kenyan ward.
- Differential generation that systematically under-weights conditions common here and over-weights conditions common in the training corpus.
- Guideline recommendations that presuppose investigations, drugs or referral pathways that do not exist at a Level 4 facility.
- Drug nomenclature and brand names that map differently in our supply chain.
- Language: the substantial degradation in performance when a history is taken in Kiswahili or Dholuo and rendered into English by the clinician, and the compounding error that introduces.
- Performance on paediatric, obstetric and geriatric presentations, systematically under-represented in the evidence base for these tools.
Every one taught with a real case and a real output, and the learner asked to find the failure before they are shown it.
The curriculum
Five professional tracks, five levels, one shared foundation.
Track C — nursing and midwifery — is the largest cadre and, in my judgement, the highest-yield group in the entire programme. Track D exists because a hospital where the clinicians are trained and the administration is not will buy the wrong system and deploy it badly.
The Level 1 common core is twelve hours, identical for the surgeon and the ward nurse, and taught in mixed-cadre groups deliberately. I am insistent on this. The consultant and the nurse should sit in the same room, because the failure modes we are trying to prevent live in the space between them — and because a nurse who has been taught to challenge an AI-supported decision needs to have practised doing so in front of a consultant who has been taught to expect it.
At the top sits a twelve-month Clinical AI Fellowship, eight fellows per cohort, open competitively to any cadre. Eight is deliberately small. I would rather produce eight people who are genuinely formidable than forty who have attended something.
How I would guarantee the quality of the deliverables
This is where most training institutions are weakest, and where I would spend a disproportionate share of my attention. It is also the part that everybody agrees with in principle and quietly dismantles in practice under delivery pressure.
A few of the specifics that matter most:
Blueprint before you build. Every module must exist as a blueprint before a word of content is written: which competency it maps to, what the learner will be able to do, how that will be assessed, what the pass standard is and how it was set. A module without a blueprint is not scheduled.
Authored by a pair. A practising clinician in the relevant cadre and an instructional designer. Neither writes alone. The clinician alone produces something accurate and unteachable; the designer alone produces something teachable and wrong. Anything unreviewed for eighteen months is automatically withdrawn from the catalogue — automatically, not on someone's judgement.
Nothing goes live unpiloted, with think-aloud observation, and pilot data goes to the Office of Quality, not to the authors. Modules that fail pilot are rebuilt, not launched with a note about improvements to follow.
Psychometrics, not vibes. Item analysis after every sitting — difficulty, discrimination, distractor analysis. Items with negative discrimination pulled immediately and affected candidates' scores recalculated. Standard setting by modified Angoff panel for knowledge assessments, borderline regression for performance assessments. Never an arbitrary 50%, and never a pass rate decided in advance. Reliability reported and published.
The assessment blueprint
The AI-OSCE deserves specific description, because I believe it would be the first assessment instrument of its kind. A candidate enters a simulated consultation with a standardised patient. They have access to an AI system in the sandbox. Unknown to them, some stations seed a clinical error into the AI's output. They are scored across four domains: appropriate delegation; quality of description; detection and correction of error; and documentation and disclosure.
The error-detection domain is a conjunctive requirement — you cannot compensate for failing it with strong performance elsewhere, in exactly the way that a candidate cannot compensate for a fatal drug error in a conventional OSCE with excellent communication skills.
Measuring what actually matters
Kirkpatrick Level 1 (reaction) we collect and largely ignore. The two that matter:
Level 3 — behaviour. At three and twelve months post-training: workplace-based assessment by a trained observer; chart audit for documentation of AI-assisted decisions; and, with consent and appropriate governance, sandbox interaction logs showing whether the independent-impression rule survived contact with real work. My working hypothesis — which I would want tested and would not be surprised to see refuted — is that the independent-impression discipline decays fastest and needs the earliest booster.
Level 4 — results. Facility-level indicators agreed in advance: documentation completeness, appropriate investigation rates, time-to-escalation for deteriorating patients, and incidents in which AI contributed to harm. Where we can run a stepped-wedge design across facilities, we should. Where we cannot, we should report the limitation honestly rather than implying causation from a before-and-after chart.
Independence, structurally
Four rules, in the founding instruments, where changing them requires a board resolution and a public explanation:
- No vendor funds curriculum development for content that concerns their own products. Unrestricted educational grants declared publicly, with amounts.
- No staff member holds equity in a company whose products the Institute evaluates. Declared annually, published.
- The Institute retains the right to publish evaluation findings regardless of outcome. Non-negotiable in any partnership agreement. If a partner will not accept it, there is no partnership.
- The teaching platform is model-agnostic by architecture, so the Institute can never become dependent on a single supplier's continued goodwill.
Plus an external examiner — a senior clinician-educator from outside Kenya, fixed non-renewable term, reporting to the board rather than to me — and international peer review of the whole institution every three years, published in full.
The team
I cannot build this alone, and I would not want to.
Why I am uniquely positioned to lead this
I want to state this directly, because it is relevant to whether this document should be taken seriously and because false modesty helps nobody.
I am uniquely positioned to build and deploy this institution. That is not a claim about being cleverer than anyone else. It is a claim about an unusual convergence of five disciplines in one career: medicine, surgery, engineering, information technology and artificial intelligence.
Medicine. I have to be able to sit with a physician and argue about a differential, and be credible. A curriculum for clinicians written by someone who has not carried clinical responsibility will be subtly wrong in ways clinicians detect immediately and then quietly disregard.
Surgery. Track B is not Track A with different examples. The decision architecture of an operating list is different from that of an outpatient clinic — the time constants are different, the reversibility is different, the relationship between information and action is different. Somebody who has stood at a table has to design that track.
Engineering. The sandbox, the seeded-error infrastructure, the logging architecture and the simulation systems are engineering problems. I need to specify them properly, judge whether a proposed architecture is sound, and know when a developer is telling me something is impossible when they mean it is inconvenient.
Information technology. Integration, interoperability, security, data governance under the Digital Health Act, and the practical realities of connectivity at a Level 4 facility. This is the layer where most well-designed health technology programmes in this region actually die.
Artificial intelligence. I have to read a model card, understand an evaluation methodology, judge a validation claim, and know precisely how these systems fail — not by analogy, but mechanically. Without this, the Institute becomes a consumer of vendor claims rather than an evaluator of them, which is the failure mode I would least be able to forgive.
Very few people anywhere hold all five, and I am not aware of anyone in this region who does. It is that combination — not any one of them in isolation — that uniquely positions me to build this institution and to deploy it.
But it also means I know exactly what I am not. I am not a psychometrician, I am not an instructional designer, and I am not a health economist. The team is built around what I cannot do.
The founding nine
| # | Post | Qualifications | Why first |
|---|---|---|---|
| 1 | Director / Chief Executive | Registered medical practitioner with surgical and technical background | Someone has to hold the whole design |
| 2 | Director of Curriculum and Pedagogy | Doctorate or master's in health professions education; competency-based curriculum design | The single most important hire — the pedagogy is the product |
| 3 | Head of Assessment (psychometrician) | Master's or doctorate in psychometrics or educational measurement; hands-on item analysis and standard-setting | Without this post the certification is worthless and we would not know it |
| 4 | Head of Engineering | CS or software engineering degree; 8+ years senior; healthcare systems and security | The sandbox must exist before the curriculum can be taught as designed |
| 5 | Clinical Lead — Medicine | Consultant physician, 7+ years post-registration, teaching experience | Track A design |
| 6 | Clinical Lead — Surgery | Consultant surgeon or anaesthetist, 7+ years post-registration | Track B design |
| 7 | Clinical Lead — Nursing and Midwifery | Senior nurse, master's-level, current or very recent practice | Track C — the largest cadre |
| 8 | Head of Quality and Evaluation | MPH, epidemiology or evaluation; independent-minded by disposition | Must be appointed at founding, not retrofitted. Reports to the board |
| 9 | Operations and Partnerships Manager | Senior programme management in health or education; regulatory navigation | Everything above fails without someone running it |
Steady state — seventy posts
| Unit | Posts | Composition |
|---|---|---|
| Executive | 1 | Director / CEO |
| Office of Quality and Evaluation (reports to the board) | 6 | Head (1); evaluation officers (2); data analyst (1); observation and audit officers (2) |
| Ethics and Data Governance | 3 | Ethics and governance lead (1); data protection officer (1, certified); research ethics coordinator (1) |
| Curriculum and Pedagogy | 14 | Director (1); instructional designers (4); clinical content leads (6); assessment psychometrician (1); medical editors and translators (3, incl. Kiswahili) |
| Faculty and Delivery | 18 | Track leads (5); certified instructors (10); simulation faculty (2); programme manager (1) |
| Engineering and Platform | 13 | Head (1); full-stack developers (5); ML and evaluation engineers (3); data engineers (2); DevSecOps (1); QA (1) |
| Simulation and Clinical Labs | 7 | Sim centre director (1); technicians (3); standardised-patient lead (1); clinical skills tutors (2) |
| Operations and Registry | 8 | Registrar and records (2); finance and procurement (2); partnerships (1); communications (1); M&E (1); admin (1) |
| Total core establishment | 70 |
Plus, not counted in the establishment: clinical champions (two per participating hospital, ~0.2 FTE sessional — without a champion on site, transfer to practice collapses within a month); visiting faculty; fellows who teach as they learn; and the patient and public panel.
How I would recruit
Hire clinicians who can teach over teachers who can clinic. Credibility with the target audience is not recoverable once lost. A nurse learner will discount a nursing module written by someone who has not been on a ward in six years, and they will be right to.
Grow the faculty from the graduates. The instructor track exists precisely so that by year three the majority of teaching is done by people the Institute trained. It is the only route to scale that does not degrade quality.
Recruit the psychometrician early and pay properly for them. Scarce skill in the region. The temptation will be to defer the post and let a clinician "handle assessment". That decision would hollow out the certification before anyone noticed.
Deliberately recruit sceptics. At least two of the clinical content leads should be people who are publicly unconvinced about clinical AI. A curriculum written entirely by enthusiasts will teach enthusiasm — and enthusiasm is the specific failure mode we are trying to prevent.
Sequence
Phase 0 — Founding (months 0–6, 9 FTE). Legal form and board. Founding nine recruited. Competency standards drafted and taken to the councils and faculties for negotiation. This phase is mostly conversation, and skipping it produces an institution nobody recognises.
Phase 1 — Prove it (months 6–15, 26 FTE). Common core written, peer-reviewed and piloted. Sandbox and de-identified case corpus built. First instructor cohort certified. A single-hospital pilot with 120 learners across three cadres. Independent evaluation published — including whatever it shows.
Phase 2 — Scale (months 15–27, 50 FTE). Tracks A–C at Levels 2 and 3 in full delivery. Tracks D and E launched. Simulation centre and AI-OSCE operational. Mobile delivery to county facilities begins. Council CPD accreditation secured.
Phase 3 — Institutionalise (months 27–36, 70 FTE). First fellowship cohort. Regional satellites. East African Community faculty exchange opens. First annual public outcomes report.
A note on the mobile unit, because it is the thing most likely to be cut: a vehicle, a generator, a satellite uplink, a set of laptops and two instructors. A substantial fraction of the workforce we most need to reach works at Level 3 and Level 4 facilities they cannot leave for a week. If the Institute becomes a thing that happens in Nairobi to people who can afford to come to Nairobi, it will have failed at its actual purpose while appearing to succeed. The metric I would watch hardest is not enrolment. It is the geographic and cadre distribution of enrolment.
What could go wrong
I would rather name these than have them named for me.
We teach enthusiasm and produce automation bias. The most likely failure and the most damaging. Mitigated by making the discernment core the largest single component of the curriculum, error-catch rate a conjunctive pass requirement, and unassisted performance a longitudinally measured outcome.
Certification becomes a box-tick. Pressure to raise throughput will arrive in year two, dressed as equity of access. Mitigated by externally set and published pass standards, an external examiner reporting to the board, and pass rates published by cohort. If a pass rate rises, someone has to explain why in public.
We become a vendor's training department. Mitigated by the four independence rules, in the founding instruments.
We serve Nairobi and call it national. Mitigated by enrolment reported by county and facility level, published quarterly, and the mobile unit funded from the core rather than from project money that can evaporate.
The technology outruns the curriculum. Mitigated by teaching judgement rather than tools, mandatory eighteen-month content review with automatic withdrawal, and a horizon-scanning function.
Trained clinicians leave. A real risk: this training makes people more employable internationally. Mitigated by a return-of-service expectation for fellows — and by the honest position that a well-trained Kenyan clinician who leaves is a loss, but a well-trained Kenyan clinician who stays untrained is a worse one. I would rather train people who might leave than protect ourselves by keeping them ignorant.
What success looks like
Five years from founding, I would want to be able to state the following publicly, with data:
- More than 15,000 clinicians across all cadres hold at least the Foundation certificate, and enrolment by county tracks workforce distribution rather than proximity to Nairobi.
- More than 2,000 hold a Practitioner certificate with a countersigned workplace logbook.
- The majority of teaching is delivered by instructors the Institute itself trained.
- Error-catch rates in the AI-OSCE have been reported for five consecutive cohorts, and the instrument's reliability has been published.
- At least one adequately-powered study, conducted here, has reported the effect of this training on a patient-relevant outcome — published whatever it found.
- The professional councils recognise the awards for CPD, and at least one has incorporated clinical AI competence into its standards.
- Clinicians from other countries are coming to Kenya to learn how this was done.
The seventh is the one I care about most. Every health system in the world is going to have to solve this problem. Most are currently solving it badly — piecemeal, vendor-led, unevaluated, and without the courage to assess whether their clinicians can actually detect an error. If Kenya solves it properly, in the open, with published outcomes and a curriculum released under a licence that permits others to adapt it, then this country will have done something for the whole of humanity that no wealthier system managed to do first.
Closing
I began by saying that AI is another arrow in the clinician's quiver. Let me be precise about what that commits me to.
It commits me to the position that the tool is not the point. The patient is the point. Every design decision here — the non-delegable list, the independent-impression rule, the seeded errors, the conjunctive pass requirement on error detection, the refusal to issue attendance certificates, the publication of unflattering results — follows from a single conviction: that a clinician's obligation to the person in front of them does not change because a new tool has arrived, and that our job as educators is to make sure the tool serves that obligation rather than quietly displacing it.
A doctor who cannot use these systems is, today, offering less than they could. A doctor who uses them without discernment is offering something worse. The distance between those two states is a curriculum, an assessment, and an institution willing to hold a standard.
I would like to build it here.
⬇ Download PDF (376 KB, 44 pages)
⬇ Download slide deck (14 MB, 15 slides)
Sources and attribution
The full reference list — seventy-one sources with links, covering Kenyan law and policy, health workforce data, WHO guidance, the automation-bias and clinical-AI evidence base, AI competency frameworks in medical education, comparator programmes, and assessment theory — is in the PDF. The principal sources for the argument above:
- The AI Fluency Framework — Dakan, R. and Feller, J. — aifluencyframework.org and the Practical Summary Document. Course series on Anthropic Academy: AI Fluency: Framework & Foundations, Teaching AI Fluency, AI Fluency for pK–12 Educators, AI Fluency for Students, Claude 101.
- Automation bias — Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial, NEJM AI, 2025.
- The Kenyan evidence — OpenAI and Penda Health, Pioneering an AI clinical copilot; the underlying study, AI-based Clinical Decision Support for Primary Care: A Real-World Study; and the critical reading in STAT News.
- Kenyan law and policy — Digital Health Act 2023; CIPESA's analysis; Kenya National AI Strategy 2025–2030.
- CPD requirements — KMPDC CPD compliance; Nursing Council of Kenya online services.
- Health workforce — Investing in the health workforce in Kenya; WHO AFRO.
- AI competencies in medical education — AAMC, Artificial Intelligence Competencies Across the Learning Continuum; WHO, Ethics and governance of AI for health: guidance on large multi-modal models.
The Clinical 4Ds are an adaptation of the AI Fluency Framework by Prof. Rick Dakan (Ringling College of Art and Design) and Prof. Joseph Feller (Cork University Business School, University College Cork), elaborated into an open course series in partnership with Anthropic PBC with support from Ireland's Higher Education Authority. The open course materials are released under CC BY-NC-SA 4.0; any curriculum derived from them carries that licence forward. The Practical Summary Document is separately released under CC BY-NC-ND 4.0 and is cited rather than adapted.
In the spirit of the framework's own Diligence competency: this document was drafted with AI assistance. The argument, the design decisions, the team composition and the pedagogical commitments are mine, and I take full responsibility for the accuracy of its contents.
