AI in healthcare
OpenEvidence's Darwin: A Medical AI Scored 100%, Then Got Locked Away
OpenEvidence's Darwin is the first medical AI to score 100% on MedQA, yet the company is withholding it over biosecurity risk. The claim, and its caveats.
When a machine passes the doctor's exam without a single mistake, and its maker decides that is a problem#
On 3 September 2026, a company few patients have heard of, but that a reported 40% of American doctors open on a typical working day, said it had built a medical AI that answered every question correctly on the field's standard exam. In the same announcement, it said almost nobody would be allowed to use that model. OpenEvidence calls it Darwin, and the decision to hold it back is arguably more revealing than the score itself.
That is the tension worth sitting with. We are used to AI launches that beg for attention. Here is one where the developer built its most capable system and then, in effect, put it behind glass. Why a medical AI firm would do that tells you something about where clinical AI has arrived, and about the anxieties now travelling alongside it.
What OpenEvidence actually released#
OpenEvidence is a Miami-based startup that runs an AI-powered medical search engine for doctors. Think of it as a search tool that reads the medical literature and answers a clinician's specific question with cited evidence, rather than handing back a page of links. On 3 September the company launched four new models at once, naming each after a founder of modern medicine.
Three are available now, free to any verified clinician. Osler, named after the physician William Osler, is the quick one, returning an answer in roughly five seconds, and it becomes the platform's default. Sackett, after David Sackett, the father of evidence-based medicine, searches more deeply and takes about 30 seconds. Snow, after John Snow, the doctor who traced an 1854 London cholera outbreak to a single water pump, is the slowest and most thorough, spending up to five minutes on the literature before writing what amounts to a short report. The company frames the difference simply: every model is held to the same accuracy standard, and what varies is time, meaning how long the model thinks and how far it searches.
The fourth model, Darwin, is the headline and the puzzle. OpenEvidence describes it as the most advanced medical AI in the world and says it is the first system ever to score a perfect result on MedQA, a widely used benchmark built from United States Medical Licensing Examination-style questions. On the company's evaluation set Darwin answered all 660 questions correctly; by comparison, it reports Claude Fable 5 at 99.7%, GPT-5.6 Sol at 99.1% and Gemini 3.7 Flash at 99.2%. Darwin also came out ahead on three harder tests, scoring 72.8% on MedXpertQA, 82.7% on HealthBench Professional and 87.2% on NOHARM.
And yet Darwin is available by application only, currently limited to a rare-disease charity, academic researchers and vetted collaborators. OpenEvidence's stated reason is dual-use risk: a model that can reason at the frontier of virology, immunology and human genetics could, in the wrong hands, help with work that is tightly governed, such as bioweapons-relevant research or gene editing outside mainstream scientific oversight. The company says Darwin's abilities will flow into the three public models only as its safeguards are checked.
A short primer: benchmarks, agents and "dual use"#
Three things are worth understanding before the rest of the story makes sense.
The first is the benchmark. MedQA is essentially a medical multiple-choice exam for machines, drawn from professional board questions. It has become the yardstick the industry quotes because it is standardised and easy to score. The catch, which matters later, is that a multiple-choice test does not resemble real clinical work, where a patient's story is messy, the right question is not printed as option C, and the job is to gather information rather than pick from four choices.
The second is the idea of an AI "agent". OpenEvidence's larger ambition, set out by founder Daniel Nadler earlier in the year, is what it calls "medical superintelligence": not one model that knows everything, but a coordinated system of specialist models, an oncology agent working next to a genetics one and a cardiology one, together behaving like a virtual tumour board a doctor can carry in a coat pocket. The Darwin launch arrived with news that the company is building exactly such an oncology sub-agent with a leading cancer centre, on top of existing licences to cancer-treatment guidelines from the NCCN and ASCO.
The third is dual use, the awkward fact that the same capability can heal or harm. A model fluent enough in immunology and genetics to crack a hard rare-disease case is, by the same token, fluent enough to be misused. That symmetry is the reason OpenEvidence gives for keeping Darwin on a leash.
Why this matters#
The immediate significance is not that a chatbot passed a test. Frontier models were already scoring in the high nineties on MedQA, so a perfect run is an increment, not a leap. What matters is the pairing of that result with a voluntary limit on access.
For the AI-safety conversation, this is one of the first times a clinical-AI developer has treated a medical model as a possible biosecurity concern and acted on it before any regulator required it to. Most of the dual-use debate has centred on general-purpose models from the big labs. Seeing a healthcare-specific company reach the same conclusion, that medical reasoning at the frontier shades into hazardous knowledge, widens the debate considerably.
For medicine, the launch sharpens a question clinicians already live with. OpenEvidence is not a fringe tool. It is used by more than a million verified US clinicians and is free at the point of use, funded by advertising and by content deals with journals including the NEJM Group and the JAMA Network. When a product at that scale changes its default model, the quality of everyday clinical answers shifts for a large slice of the profession overnight. Nadler's own framing is that this could level access to specialist knowledge in rural and under-resourced areas, a real prize if it holds up in practice.
The caveats the headline number hides#
A perfect score invites scrutiny, and this one earns it.
Start with how the number was produced. Darwin did not ace the standard MedQA test as most people would picture it. OpenEvidence physicians re-annotated the benchmark, removing questions they judged ambiguous, incomplete or mislabelled, which cut the original 1,273-question split to a cleaned set of 660. The perfect score is on that curated set, scored by the company itself. Cleaning a noisy benchmark is defensible, but this is a company-run evaluation, not an independent audit, and "100%" travels a lot further than the asterisk attached to it.
Then there is the deeper problem the company itself concedes. These benchmarks test a model working alone, with no clinician in the loop, which does not match how the tools are actually used. In practice the model is an aid to a doctor's judgement, and an answer is only good if it helps that doctor decide better. Measured that way, a leaderboard tells you very little.
The most pointed challenge is recent. In June 2026, researchers from NYU Langone Health published a study in Nature Medicine that pitted specialised clinical tools, including an earlier OpenEvidence model and UpToDate's Expert AI, against general-purpose models from the big labs, across medical-knowledge questions and real clinical queries. The general models won across the board. One reading of Darwin, then, is a specialist tool reasserting itself after an uncomfortable result, which is worth cheering only once someone outside the company has checked the work. That June paper also drew a formal published rebuttal and reply in Nature Medicine over whether its benchmarks were fair, a sign the field has not settled how to measure any of this.
Finally, MedQA itself is fraying as a yardstick. Because leading models already sit between 96% and 99%, the test has little headroom left, and researchers warn of possible contamination, the chance that benchmark questions leaked into training data so a model recalls rather than reasons. A perfect score on a saturated exam is a strange trophy: impressive, and also a hint that the exam has outlived its use.
How this compares with what came before#
Set Darwin next to the milestones and the shape of the change becomes clearer. When Google's Med-PaLM cleared the MedQA passing mark in 2023, the story was that a machine could pass the medical boards at all. Three years on, passing is assumed, near-perfect scores are common, and the interesting differences have moved elsewhere: to speed, to depth of literature search, to safety, and to whether a tool actually improves a clinician's decision rather than its own test score.
What makes OpenEvidence's move different is less the model than the packaging around it. Rather than one system, it shipped a tiered family and let clinicians choose how much thinking time a question deserves, five seconds for a routine query or five minutes for a thorny differential. And rather than releasing its best model, it withheld it on safety grounds. Competitors have taken other routes: the big labs push powerful generalist assistants at everyone, while incumbents such as UpToDate and Elsevier lean on curated, human-edited authority. OpenEvidence is trying to occupy the middle, with frontier capability, clinician-verified access and a self-imposed brake, betting that doctors will trust the combination.
Key takeaways#
- OpenEvidence released four medical AI models on 3 September 2026. Three of them, Osler, Sackett and Snow, are free to verified clinicians and differ mainly in how long they take to answer.
- The fourth, Darwin, is the first model the company says has scored a perfect result on the MedQA benchmark, but only on a curated 660-question set that OpenEvidence re-annotated and scored itself.
- Darwin is being withheld from general release over dual-use concerns: the risk that frontier-level reasoning in virology, immunology and genetics could aid bioweapons or unsupervised gene editing.
- Benchmarks test models in isolation. A June 2026 Nature Medicine study found general-purpose models beat specialised clinical tools on real clinical queries, and the field still disagrees on how to measure clinical AI at all.
- The lasting significance is the precedent: a clinical-AI company treating its own medical model as a biosecurity matter and restricting it voluntarily, before any regulator asked.
Frequently asked questions#
What is OpenEvidence? It is a medical search engine and AI assistant for doctors, sometimes described as a "ChatGPT for doctors". Clinicians ask a question about a patient and it returns an answer grounded in cited medical literature. It is free to verified clinicians and funded by advertising and journal content partnerships.
What is Darwin, and can I use it? Darwin is OpenEvidence's most advanced model, currently in research preview. It is not generally available. Access is by application and limited to institutional partners such as the National Organization for Rare Disorders, academic researchers and vetted collaborators.
What is MedQA? A benchmark made of United States medical licensing-style multiple-choice questions, used to measure a model's medical knowledge. It is standardised and easy to score, but it does not capture the ambiguity of real patient encounters.
Did Darwin really score 100%? On OpenEvidence's own curated set of 660 questions, yes. That set was produced by removing questions the company's physicians judged flawed from the original 1,273. It is a company-run evaluation rather than an independent one.
Why hold back a model that helps doctors? Because the same reasoning that solves a hard rare-disease case could, in principle, assist dangerous work in virology or genetics. OpenEvidence says it will feed Darwin's abilities into its public models only as safety checks are validated.
Is this AI going to replace my doctor? No. These tools are positioned as aids to clinical judgement, not substitutes. The company and outside experts stress that the goal is better human-plus-AI decisions, and that benchmark scores say little about real-world care.
Should patients trust answers that came from it? Trust rests with the clinician using the tool, who is responsible for the decision. Independent, real-world evaluation, not leaderboard scores, is what will show whether these systems improve care.
Glossary#
Benchmark: A standard test used to compare AI systems. In medicine, examples include MedQA and HealthBench.
MedQA: A benchmark of US medical licensing-style multiple-choice questions used to gauge a model's medical knowledge.
Dual-use: Knowledge or technology that can serve both beneficial and harmful ends. Here, medical reasoning that could also aid bioweapons or unsupervised gene editing.
AI agent: A model set up to carry out tasks and, often, to specialise. OpenEvidence's plan chains specialist agents, such as oncology and genetics, into one system.
Medical superintelligence: OpenEvidence's term for a coordinated ensemble of specialist medical AI agents meant to act like a full multi-speciality care team.
Large language model (LLM): An AI trained on vast amounts of text to generate and reason over language. It is the technology underlying tools like OpenEvidence and general chatbots.
Benchmark contamination: When test questions leak into a model's training data, so strong scores may reflect memorisation rather than reasoning.
Clinical decision support: Software that helps clinicians make decisions by supplying relevant evidence, guidance or alerts at the point of care.
References#
- OpenEvidence. Introducing the OpenEvidence Model Family. Company announcement, 3 September 2026.
- Fierce Healthcare. OpenEvidence deepens oncology push to build specialized AI agents, launches new AI model family. 3 September 2026.
- Unite.AI. OpenEvidence Launches Medical AI Model Family With Darwin Preview. 3 September 2026.
- TechTarget / Healthtech Analytics. OpenEvidence launches 4 medical AI models. 4 September 2026.
- Nature Medicine. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. June 2026.
- Nature Medicine. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Correspondence and reply, 2026.
- STAT. Clinical chatbots are taking medicine by storm. Should doctors trust them?. 29 July 2026.
- CNBC. OpenEvidence, the 'ChatGPT for doctors,' doubles valuation to $12 billion. 21 January 2026.
- Fierce Healthcare. JAMA signs multi-year deal with OpenEvidence.
- MedXpertQA. Benchmarking Expert-Level Medical Reasoning and Understanding. arXiv preprint (not peer reviewed).