The Patient Who Waited Twenty Years#

Imagine living with symptoms nobody can name. You see a paediatrician, then a neurologist, then a geneticist. Tests come back normal, or ambiguous, or unhelpful. Years pass. Then someone feeds your records to an AI system, and it suggests a new lead that a laboratory later confirms.

That is not a thought experiment. In a study from Boston Children's Hospital, Harvard and OpenAI, researchers re-examined 376 previously unsolved rare-disease cases with a reasoning model and, after expert review and certified laboratory confirmation, reached new diagnoses in 18 of them. One was a patient who had gone almost 20 years without an answer (OpenAI).

The headline is tempting: AI diagnoses rare disease in minutes. The reality is more interesting. Nobody in that study measured how much time was saved, and the authors said the model proposes hypotheses rather than diagnoses (OpenAI). So the honest question is not "can AI diagnose?" but "what happens when a doctor and an AI try to work together?"

This post walks through the best current evidence on AI diagnosis, including the successes, the stumbles and the uncomfortable finding that a clinician plus an AI is not always better than the AI alone.

Why Rare Diseases Are So Hard to Diagnose#

A rare disease is, by definition, uncommon in any single population. Collectively they are not: rare diseases affect more than 300 million people worldwide, and many patients endure a "diagnostic odyssey" lasting more than five years, with repeated referrals, misdiagnoses and unnecessary interventions (Zhao et al., Nature).

The reason is mostly arithmetic. There are thousands of rare conditions, and a typical doctor may meet any given one only once in a career, if ever. One research group estimated that around 70% of people seeking a rare-disease diagnosis remain undiagnosed (Alsentzer et al., npj Digital Medicine).

A few terms are worth defining before we go further.

Phenotype means a person's observable traits and symptoms, such as seizures, a distinctive facial feature or muscle weakness. Doctors often record these using the Human Phenotype Ontology (HPO), a standard vocabulary that lets computers compare patients.

Differential diagnosis is the ranked list of conditions a clinician considers before settling on one.

Large language models (LLMs) are the AI systems behind chatbots. They predict text, but they can be wrapped in tools that let them search databases, read genetic reports and cite sources.

Agentic systems are collections of such models that split a problem into steps, call external tools and check one another's work, a little like a hospital multidisciplinary team in software form.

Recall@1 is a simple scoring rule: how often is the correct diagnosis the system's top-ranked answer? Recall@5 counts it as a success if the right answer appears anywhere in the top five.

Finally, the word that keeps this field honest: benchmark. A benchmark is a set of test cases with known answers. Strong benchmark scores do not automatically translate into better care on a busy ward.

What the Strongest Rare-Disease Systems Can Do Now#

The most prominent recent example is DeepRare, published in Nature. It is a multi-agent system that integrates more than 40 specialised tools and knowledge sources, accepts free-text notes, HPO terms and genetic test results, and produces ranked diagnoses with reasoning linked to verifiable evidence (Zhao et al., Nature).

It was evaluated across nine datasets from Asia, North America and Europe, covering 14 medical specialties and 2,919 diseases. On HPO-based tasks it reached an average Recall@1 of 57.18%, 23.79 percentage points ahead of the next best method (Zhao et al., Nature). When genetic data were added, it scored 69.1% against 55.9% for Exomiser, a widely used genetics tool, on 168 cases (Zhao et al., Nature).

Two features matter beyond the raw scores. First, the reasoning is traceable. Expert reviewers agreed with 95.4% of the system's reasoning chains (Zhao et al., Nature). A black-box answer is hard for a clinician to trust, whereas a chain of evidence can be checked. Second, there was a head-to-head test against people. According to a summary of the paper, five physicians with at least ten years of rare-disease experience, allowed to use search engines but not AI, achieved Recall@1 of 54.6% and Recall@5 of 65.6%, against 64.4% and 78.5% for DeepRare on the same cases (Lifespan.io). Nature's accompanying news piece framed the result under the title "AI succeeds in diagnosing rare diseases" (Lassmann, Nature).

DeepRare versus experienced physicians on identical rare-disease cases Grouped bar chart. Recall at 1: physicians 54.6 per cent, DeepRare 64.4 per cent. Recall at 5: physicians 65.6 per cent, DeepRare 78.5 per cent. Top-ranked and top-five accuracy on identical cases (%) 0 20 40 60 80 54.6 64.4 65.6 78.5 Recall@1 (correct answer ranked first) Recall@5 (correct answer in top five) Experienced physicians (n = 5) DeepRare

Figure: DeepRare against five experienced physicians on identical cases, as reported in a summary of the Nature paper (Lifespan.io).

It is an impressive result, and it still has limits. The comparison involved a small panel of five doctors. Some newer models were not tested, since the strongest ChatGPT version used was GPT-4o (Lifespan.io). And scoring well on retrospective case files is not the same as navigating a real clinic, with missing notes and a worried family.

Solving Cold Cases: The Boston Children's Study#

If DeepRare is the benchmark winner, the Boston Children's Hospital work is the reality check. Published in NEJM AI in June 2026, it used OpenAI's o3 Deep Research model to reanalyse de-identified clinical and genomic data from cases that specialists had previously failed to solve (OpenAI).

Yields varied by cohort, which is the useful part. The table below summarises them.

Patient groupUnsolved cases reviewedNew diagnosesYield
Neurodevelopmental1001010.0%
Neuromuscular6146.6%
Early psychosis15213.3%
Sudden unexpected death20021.0%
All cohorts376184.8%

Source: OpenAI summary of the NEJM AI study.

A 4.8% overall yield sounds modest until you remember these were cases already considered stuck. For the families affected, an answer can change management, open up support groups or end years of uncertainty. The model also proposed novel two-gene (digenic) explanations and testable hypotheses (OpenAI).

The caveats were frank. The design was retrospective, reviewers were not blinded, structural variants and repeat expansions were not evaluated, and time or clinician effort saved was not measured (OpenAI). Every diagnosis also needed expert review and confirmation in an accredited laboratory (OpenAI). The AI's role was to generate leads, and humans did the verifying. That division of labour is probably the realistic template for the next few years.

Do Doctors Actually Get Better With AI?#

Here is where the story gets less tidy. The cleanest way to test whether AI helps doctors is a randomised controlled trial, in which similar clinicians are randomly assigned to work with or without the tool.

The first major trial, published in JAMA Network Open in 2024, gave 50 physicians six clinical vignettes. Those with access to an LLM scored a median 76% against 74% for those using conventional resources, a difference too small to count as significant. The LLM on its own scored 16 percentage points higher than the conventional group (Goh et al., JAMA Network Open). In plain terms: the AI was good, the doctors were good, and together they were not much better than the doctors alone, probably because the doctors had not been taught how to use it.

Later trials suggest training matters. In a Nature Health trial of 58 Pakistani physicians who completed a 20-hour AI-literacy course, those using an LLM scored 71.4% against 42.6% with conventional resources, without taking longer per case (Qazi et al., Nature Health). Even so, the LLM alone beat the LLM-assisted physicians by 11.5 percentage points (Qazi et al., Nature Health).

Workflow design also appears to matter. An npj Digital Medicine trial of 70 clinicians compared AI as a first opinion, AI as a second opinion and conventional resources. Both AI workflows lifted accuracy (85% and 82%) over conventional resources (75%), landing close to the AI-alone figure of 90% (Everett et al., npj Digital Medicine). A study in critical care found residents' top-diagnosis accuracy rose from 27% to 58% with AI help, with diagnostic time roughly halved (Wu et al., Critical Care).

And a widely discussed paper in Science reported that an LLM outperformed physician baselines across five experiments and in a real-world emergency-room second-opinion study, while stressing the "urgent need for prospective trials" (Brodeur et al., Science).

A pooled analysis keeps the champagne corked. A 2026 meta-analysis of human-and-LLM collaboration found the evidence "preliminary yet highly uncertain and context-dependent", with wide prediction intervals and high error rates in AI-generated documentation (Wang et al., npj Digital Medicine). Another systematic review found most included studies at high risk of bias, with top-diagnosis accuracy for the best model ranging from 25% to 97.8% (Shan et al., JMIR Medical Informatics).

ai diagnosed a rare disease in minutes should your doctor use it 2

When the AI Is Wrong: Automation Bias and Other Risks#

The least comfortable finding concerns what happens when the machine makes a mistake. Automation bias is the tendency to over-trust an automated suggestion, even when your own judgement should have caught the error. Think of following a satnav into a field.

One randomised preprint (not peer reviewed) tested exactly this with AI-trained physicians. Doctors given error-free AI advice scored 84.9% on diagnostic reasoning, while those shown deliberately flawed advice in three of six cases scored 73.3%, a drop of 14 percentage points (Ali et al., preprint, not peer reviewed). Because it is a preprint, it should be treated as a signal rather than settled fact. Still, it points at a design problem: tools that present one confident answer can quietly narrow a clinician's thinking.

This is one reason traceable reasoning, as in DeepRare, matters so much. A ranked list with visible evidence invites checking. A single polished paragraph invites trust.

Other risks are worth naming, in brief. Training data may under-represent some populations, so performance can vary between groups. Benchmarks often use tidy case reports, whereas real notes are messy. Privacy matters too, because rare-disease patients are, almost by definition, easy to re-identify. And the infrastructure matters: the Boston study stressed that genetic counselling and confirmatory testing must exist for an AI lead to become a diagnosis (OpenAI).

So, Should Your Doctor Be Using It?#

On the evidence so far, a fair answer is "yes, as a second pair of eyes, with conditions."

The strongest case is for the hard, stuck cases. Rare-disease tools are not replacing a clinician's examination. They are searching a literature and a catalogue of genetic variants that no human can hold in their head. For a patient on year six of an odyssey, an extra list of possibilities is cheap and potentially life-changing.

The conditions are fairly clear. Clinicians need training in how these systems work and fail, which the trials suggest makes a large difference. Tools should show their evidence and uncertainty rather than a lone verdict. Hospitals should compare doctor-plus-AI against doctor-alone in real workflows, and measure harms as well as hits. Regulators and procurement teams should ask for prospective evidence, not only benchmark scores.

Meanwhile, patients can reasonably ask a simple question at an appointment: "If we are stuck, is there a decision-support tool or a genetics service we could try?" It is a polite, specific request, and it keeps the human in charge.

Evidence at a Glance#

StudyDesignKey result
DeepRare, NatureBenchmarks across nine datasetsRecall@1 57.18% on HPO tasks; 95.4% expert agreement on reasoning
Boston Children's / OpenAI, NEJM AIRetrospective reanalysis of 376 unsolved cases18 new diagnoses (4.8%) after expert and laboratory confirmation
Goh et al., JAMA Netw OpenRandomised trial, 50 physiciansNo significant gain for doctors with an LLM; LLM alone scored higher
Qazi et al., Nature HealthRandomised trial, 58 trained physicians71.4% with LLM vs 42.6% without
Everett et al., npj Digit MedRandomised trial, 70 clinicians85% / 82% with AI workflows vs 75% conventional
Ali et al., preprint (not peer reviewed)Randomised trial, 44 physiciansFlawed AI advice cut accuracy by about 14 points

Frequently Asked Questions#

Did an AI really diagnose a rare disease in minutes? Not in a way the studies measured. The Boston Children's study reported 18 new diagnostic leads confirmed after expert review, but it did not measure time saved, and the model generates hypotheses rather than final diagnoses (OpenAI).

How accurate are AI systems for rare diseases? It depends on the task and test. DeepRare reached an average Recall@1 of 57.18% on HPO-based tasks across nine datasets, meaning the right answer was ranked first just over half the time (Zhao et al., Nature). That is strong for a field this hard, and still means many wrong first answers.

Does AI beat doctors at diagnosis? Sometimes, on benchmarks. DeepRare outscored five experienced physicians in a head-to-head test (Lifespan.io), and LLMs alone often outperform doctors without AI on vignettes (Goh et al., JAMA Network Open). Real-world prospective evidence is still thin (Brodeur et al., Science).

Why does a doctor using AI sometimes do no better than the AI alone? Possible reasons include limited training, how the tool is presented and how clinicians weigh its advice. Trained physicians gained substantially in one trial, but the AI alone still scored higher (Qazi et al., Nature Health).

What is automation bias? It is over-reliance on an automated suggestion. In a randomised preprint (not peer reviewed), physicians shown flawed AI advice scored about 14 percentage points lower than those shown correct advice (Ali et al., preprint, not peer reviewed).

Is AI safe to use for my own diagnosis? Clinical tools are designed for use by trained professionals, and the best-performing systems need expert review and laboratory confirmation before anything is diagnosed (OpenAI). General chatbots should not replace a medical consultation. If you are worried, discuss it with your doctor or a genetics service.

Which patients might benefit most? Those on a prolonged diagnostic odyssey. Such journeys often exceed five years (Zhao et al., Nature), so even a modest number of new leads could matter to families who have exhausted standard routes.

What needs to happen before AI diagnosis becomes routine? Researchers call for prospective, multicentre trials embedded in real workflows that prioritise safety and error metrics, with interfaces that surface uncertainty (Wang et al., npj Digital Medicine).

References#

  1. Zhao W, Wu C, Fan Y, et al. An agentic system for rare disease diagnosis with traceable reasoning. Nature (2026). https://doi.org/10.1038/s41586-025-10097-9
  2. Lassmann T. AI succeeds in diagnosing rare diseases. Nature 651, 597-598 (2026) (news). Indexed in PubMed via PubMed 41708821; DOI 10.1038/d41586-026-00290-9
  3. OpenAI, Boston Children's Hospital and Harvard. Using AI to help physicians diagnose rare genetic diseases affecting children (summary of the NEJM AI study, June 2026). https://openai.com/index/diagnose-rare-childhood-diseases/
  4. Lifespan.io. AI tool sets new standard in diagnosing rare diseases (science journalism summarising the DeepRare physician comparison). https://lifespan.io/news/ai-tool-sets-new-standard-in-diagnosing-rare-diseases/
  5. Goh E, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Network Open (2024). Link
  6. Qazi I, et al. Large language model diagnostic assistance for physicians in a lower-middle-income country: a randomized controlled trial. Nature Health (2026). Link
  7. Everett S, et al. From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis. npj Digital Medicine (2026). Link
  8. Wu X-T, et al. A large language model improves clinicians' diagnostic performance in complex critical illness cases. Critical Care (2025). Link
  9. Brodeur P, et al. Performance of a large language model on the reasoning tasks of a physician. Science (2026). Link
  10. Wang G, et al. Human-large language model collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine (2026). Link
  11. Shan G, et al. Comparing diagnostic accuracy of clinical professionals and large language models: systematic review and meta-analysis. JMIR Medical Informatics (2025). Link
  12. Alsentzer E, et al. Few shot learning for phenotype-driven diagnosis of patients with rare genetic diseases (SHEPHERD). npj Digital Medicine (2025). Link
  13. Ali A, et al. Automation bias in large language model assisted diagnostic reasoning among AI-trained physicians. Preprint (not peer reviewed). Link