DULUTH, MN MANILA, PH
Say hello

Case Study 2026

Five real clinical-AI deployments, and the line that separates proof from promise

EvaluatingClinical AI

Summary. I traced five real, currently-deployed clinical AI technologies back to their primary evidence — not vendor decks, not conference abstracts, the actual trials. Autonomous AI reads retinal photos and diagnoses diabetic retinopathy without a human ever looking at the image. AI watches the colonoscope feed and points at polyps. AI reads mammograms alongside — sometimes instead of — a second radiologist. AI reads a standard ECG and flags patients who are about to die. And a conversational AI now outperforms physicians on standardized diagnostic vignettes. Every one of these has real, reproducible evidence that it detects more of something. Only one of them — the ECG — has a randomized trial showing that detecting more of something actually changed how many patients died. That gap, between “finds more” and “saves more,” is the single question worth asking about any clinical AI pitch, including the next one you see.

The line that matters

Every technology below clears the first bar: a real peer-reviewed trial, often several, showing it out-detects the status quo on some measurable target — an adenoma, a low ejection fraction, a cancer on a mammogram, the right diagnosis on a vignette. That bar is not nothing; it used to be the whole argument for adopting a tool. But detecting more of a proxy is not the same claim as preventing more of the bad outcome, and for four of these five technologies, the trial that would prove the second claim either hasn’t been run, hasn’t reported, or came back short of significance. I’m not against the technologies below — I use one of the underlying categories (validated decision support) every shift. I am against collapsing “detects more” into “helps more” without checking whether anyone actually measured the second thing.

Autonomous retinal screening: real detection, a real-world specificity gap

IDx-DR (now LumineticsCore) became the first autonomous AI diagnostic device ever authorized by the FDA in April 2018 — no ophthalmologist reviews the image; the software’s read is the diagnosis. Its pivotal trial was real: 900 patients across 10 primary care sites, sensitivity 87.2%, specificity 90.7%, both clearing FDA’s pre-specified bar (Abràmoff et al., npj Digital Medicine 2018). A second device, EyeArt, cleared in 2020 on similarly strong numbers (sensitivity 96–97%, specificity 88–90%, n=942, Ipp et al., JAMA Network Open 2021).

What doesn’t travel from the trial into the clinic as cleanly: a 2025 multi-institution review spanning Washington University, Johns Hopkins, Stanford, Mayo Clinic, and others found real-world image gradability running 49–75%, against 96.1% in the pivotal trial, and site-level specificity as low as 60% (Teng et al., Ophthalmology Science 2025). An independent Swiss validation found specificity as low as 78.4% and a positive predictive value as low as 0% in one severity category — the device was overcalling disease that specialists didn’t confirm (Riotto et al., J Clin Med 2024). The genuinely good news is on the access side, not the accuracy side: Johns Hopkins data show AI-screened clinics posted a 7.6-point larger rise in annual DR-testing adherence than non-AI clinics, with the gain concentrated in Black patients (+12.2 points vs. −0.6 at non-AI sites) — a real, measured equity effect (npj Digital Medicine 2026).

AI-ECG: the one place a hard outcome actually moved

This is the strongest evidence in this piece, and it’s the one most physicians haven’t heard about. A Mayo Clinic-derived deep-learning model reads a standard 12-lead ECG and flags patients likely to have a low ejection fraction or to be at elevated near-term mortality risk. The foundational trial (120 primary care teams, 22,641 patients) showed the AI flag increased new low-EF diagnoses by 32% (Yao et al., Nature Medicine 2021). The finding that actually matters came in 2024: a pragmatic randomized trial of 15,965 hospitalized patients in Taiwan found the AI-ECG alert reduced 90-day all-cause mortality from 4.3% to 3.6% — hazard ratio 0.83, effect concentrated in the highest-risk-flagged patients (HR 0.69) (Lin et al., Nature Medicine 2024). A 2025 follow-up in 13,631 inpatients from the same group replicated the diagnostic-rate effect (Tsai et al., BMC Medicine 2025), and a secondary cost-effectiveness analysis of the underlying screening trial put it at $27,858/QALY — genuinely cost-effective by conventional thresholds (EAGLE trial secondary analysis, 2024).

The honest caveat: the mortality result and its replication both come from the same Taiwanese academic center. That’s real replication, not yet independent multi-country confirmation of the specific claim that this alert saves lives — and the absolute risk reduction, 0.7 percentage points, is real but small at the single-patient level. I’d still call this the most convincing clinical-AI evidence I reviewed for this piece, and it’s the one I’d want deployed first.

AI-assisted colonoscopy: strong on paper, contested at the bedside

Computer-aided polyp detection (CADe) has the deepest trial base of anything here — more than 20 randomized trials, several independent meta-analyses. The original landmark trial found adenoma detection rate (ADR) of 54.8% with AI versus 40.4% without (Repici et al., Gastroenterology 2020); pooled analyses since have consistently found the AI arm ahead by roughly 10–25% relative, though not uniformly — a 2025 meta-analysis limited to seven trials of one specific device found no significant advantage for detecting the higher-stakes advanced adenomas specifically (RR 1.01, P=0.85, PMC12616575), and a more recent 12-trial pooled analysis found the overall ADR effect itself fell short of significance (OR 1.24, 95% CI 0.98–1.58, P=0.08, PMC12524564) — the same analysis that found no included trial has ever reported colorectal-cancer-incidence outcomes (Ann Intern Med 2023).

Real-world deployment data is genuinely split, which is the part worth knowing before you read any single positive headline. A 334,200-colonoscopy VA rollout found a real gain (ADR 50.7%→54.9% at AI sites, flat at controls) — reported via Medscape’s summary of the quality-improvement dataset; I have not independently pulled the primary study, so treat the exact figures as secondary-sourced. A more rigorous Stanford pragmatic implementation — same device, real clinic, no extra fanfare — found no difference at all (ADR 40.1% vs. 41.8%, P=.44), and the authors’ own hypothesis is uncomfortable: much of the trial-measured benefit may reflect endoscopists knowing they were being studied, not the algorithm itself (Ladabaum et al., Gastroenterology 2023). A 2025 multicenter Polish study raised a further concern that should worry anyone deploying this today: after routine AI exposure, the same experienced endoscopists’ unassisted ADR fell from 28.4% to 22.4% — a real, if unreplicated, deskilling signal (Budzyń et al., Lancet Gastroenterol Hepatol 2025). And no trial or meta-analysis I found has reported the outcome that would actually settle the argument: whether any of this changes interval colorectal cancer incidence.

AI-supported mammography: a headline number that didn’t reach significance

MASAI, the Swedish randomized trial of AI-supported mammography reading, is real, large (n≈106,000), and genuinely well-run. Its interim safety analysis showed a small cancer-detection increase and cut radiologist reading workload by 44% (Lång et al., Lancet Oncology 2023). Its final results, reported in early 2026 after two years of follow-up, are the ones worth reading carefully: interval cancer rate (the outcome that actually proves AI catches cancers a human would otherwise miss) was 1.55 per 1,000 with AI versus 1.76 per 1,000 without — a favorable direction, but the proportion ratio (0.88, 95% CI 0.65–1.18, P=0.41) did not reach statistical significance. What the trial did establish, and what its own pre-specified hypothesis actually was, is non-inferiority — AI-supported reading is not worse — plus a significant sensitivity gain (80.5% vs. 73.8%, P=0.031). “12% fewer interval cancers” is how this result has often been reported; the honest version is “a favorable but statistically inconclusive trend.”

Two large real-world US and German deployments (579,583 and 463,094 exams respectively) independently found 17–22% higher cancer detection with AI — but neither is randomized, so the gain could partly reflect secular trends rather than the algorithm (ASSURE, Nature Health 2025; PRAIM, Nature Medicine 2025). A US-specific randomized trial designed to test this properly — MASAI’s double-reading design doesn’t map onto how American radiologists actually read mammograms — launched in late 2025 (PRISM, PCORI-funded, 7 academic sites) and has not yet reported a single result.

LLM diagnostic reasoning: ahead on paper, unproven at the bedside

Google DeepMind’s AMIE, a conversational diagnostic AI, has now appeared in four studies, and the distinction between them is the whole story. In a randomized crossover study, AMIE was rated superior to 20 primary care physicians on 30 of 32 axes by specialist evaluators — but both AMIE and the physicians were typing into a chat window with actors playing scripted patients, not seeing real patients (Nature 2025). Run against 302 real, published diagnostic case records from the New England Journal of Medicine, standalone AMIE hit the correct diagnosis in its top 10 guesses 59.1% of the time versus 33.6% for unassisted physicians working the same records (Nature 2025). The one study that used AMIE on real patients — a 100-patient Boston primary-care pilot, every interaction supervised by a human physician who could intervene, no control arm — remains an unreviewed preprint as of this writing.

The finding physicians should actually sit with is a negative one, and it didn’t involve AMIE at all. A 2024 randomized trial gave 50 physicians either their usual references or their usual references plus an LLM, working through real diagnostic vignettes. The LLM working alone scored 92% — significantly ahead of physicians working alone at 74%. But physicians given the same LLM scored 76%, statistically indistinguishable from the 74% they scored without it (Goh et al., JAMA Network Open 2024). The tool’s solo advantage did not transfer to the person using the tool. A separate trial found GPT-4 scored higher than residents and attendings on a validated clinical-reasoning documentation rubric — while also being flatly wrong more often than the residents were (13.8% vs. 2.8%) (Cabral et al., JAMA Intern Med 2024). Sounding more rigorous and being more correct are not the same measurement, and as of mid-2026 no randomized trial has tested an LLM used without physician supervision as a diagnostic aid on real patients.

Two more worth watching, more briefly

AI dermatology triage (DermaSensor, FDA-authorized Jan 2024, the first AI device covering all three common skin cancers) is explicitly scoped as an adjunct for lesions a clinician has already flagged suspicious, not a screen. Its pivotal trial found sensitivity 95.5% versus 83% for unassisted primary care physicians — with specificity of only 20.7%, meaning most positive flags are false alarms (npj Digital Medicine 2024). A Class 2 FDA recall for a subset of devices — potential for incorrect results or delayed referral — has been open since October 2025 (FDA recall record Z-0583-2026).

AI sepsis prediction beyond the Epic Sepsis Model — the Prenosis Sepsis ImmunoScore became the first FDA-authorized AI sepsis diagnostic in April 2024, with a genuine external-validation cohort the Epic Sepsis Model never had (NEJM AI 2024). But it’s a correlation/validation study, not a trial proving that acting on the score changes outcomes. An independent Korean validation of a different sepsis-AI tool found excellent discrimination (AUROC 0.88) alongside a positive predictive value of just 3.7% — most alerts, in that population, were false alarms (Aitrics VC-SEPS validation).

What I’d want as a hospitalist

  1. Ask the outcome question before the detection question. “Does it find more X” is the easy yes. “Does finding more X change how many patients are harmed” is the question that’s actually been answered for exactly one of the five technologies above.
  2. A real-world validation gap is not disqualifying — it’s just a fact to plan around. Retinal screening’s specificity drop, colonoscopy’s Stanford null result, and the deskilling signal are all reasons to build a workflow that assumes the trial numbers won’t fully travel, not reasons to reject the technology outright.
  3. Same three questions I ask of any decision-support tool: validated on whom, published where, and who owns the decision when the model is wrong. AI-ECG is the only technology here I’d currently answer all three questions about to my own satisfaction.

Claim-by-claim scorecard

Technology Best evidence Detects more? Changes outcomes? Verdict
Autonomous retinal screening Abràmoff, npj Digital Med 2018; Teng, Ophthalmology Sci 2025 Yes, in trial (sens 87–97%) Real-world specificity drops to 60–91%; access/adherence gains are real Verified detection, contested real-world accuracy
AI-ECG mortality alert Lin, Nat Med 2024 Yes Yes — 90-day mortality 4.3%→3.6%, HR 0.83, RCT Verified — strongest evidence of the five
AI colonoscopy (CADe) Repici 2020; Ann Intern Med meta-analysis 2023 Yes in most pooled analyses, not all No RCT has reported CRC-incidence outcomes; deskilling signal emerging Verified detection, outcome and durability unproven
AI mammography (MASAI) Lång, Lancet Oncol 2023; MASAI final results, 2026 Yes, sensitivity significant Interval-cancer reduction NOT statistically significant (P=0.41) Non-inferiority proven; superiority on the outcome that matters is not
LLM diagnostic reasoning AMIE, Nature 2025; Goh, JAMA Netw Open 2024 Yes, on vignettes/records RCT found LLM-assisted physicians did NOT significantly outperform unassisted ones Ahead on paper, unproven with a physician in the loop
AI dermatology triage DermaSensor, npj Digital Med 2024 Yes (sens 95.5% vs. 83%) Specificity 20.7%; open FDA Class 2 recall since Oct 2025 Real triage aid, real false-positive cost
AI sepsis prediction Prenosis, NEJM AI 2024; VC-SEPS validation Yes (AUROC 0.80–0.88) No interventional trial yet; PPV as low as 3.7% in one validation Accuracy proven, deployment benefit unproven

— Jeremy Tabernero, MD · More case studies · Get in touch