AI
What the VCT Benchmark Means for Pandemic Preparedness
A multimodal benchmark of 322 expert-validated questions finds OpenAI's o3 outperforms 94% of PhD virologists on troubleshooting complex lab protocols. We unpack what changed, what it doesn't prove, and why pandemic-preparedness planners should care.
A new benchmark, the Virology Capabilities Test (VCT), OpenAI's o3 scored 43.8%, surpassing 94% of PhD virologists in their sub-fields, where the human baseline was 22.1%[1]. This shift affects how AI labs, biosecurity groups, and governments view the boundary between helpful and harmful AI in life sciences.
The VCT is a 322-question, multimodal benchmark co-developed by SecureBio, the Center for AI Safety, the Federal University of ABC (Brazil), and MIT Media Lab[1]. The questions were written and peer-reviewed by dozens of PhD-level virologists, cover troubleshooting of real wet-lab protocols (plaque assays, cell culture, viral stock production, cloning, etc.), and many are deliberately search-proof — they test tacit and visual knowledge that a Google query won't surface. Each question presents an experimental scenario, often with an image, and asks what went wrong or what to do next. The team explicitly designed the test to span "fundamental, tacit, and visual knowledge essential for practical work in virology laboratories," and to be relevant to dual-use topics without crossing into the most hazardous content.
The numbers (from VCT preprint):
| Model | VCT accuracy | Expert percentile |
|---|---|---|
| Expert virologists (in own sub-field) | 22.1% | — |
| OpenAI o3 | 43.8% | 94th |
| Google Gemini 2.5 Pro | 37.6% | 81st |
| OpenAI o4-mini | 37.0% | 86th |
| OpenAI o1 | 35.4% | 89th |
| Anthropic Claude 3.7 Sonnet (Oct '24) | 33.6% | 75th |
Two findings stand out. First, every frontier model tested cleared the median human expert. Second, the LLM-over-human edge is not new: SecureBio's own writeup notes that "the first LLM to beat the median expert virologist, Gemini 1.5 Pro, was released in February 2024" — meaning publicly available models have been at or above median expert level for over two years at the time of writing, and the gap is widening[2].
OpenAI's own o1 system card [4], released earlier, had already acknowledged that "o1 can help experts with the operational planning of reproducing a known biological threat, which meets our medium risk threshold"[4]. The VCT extends that finding from a single company's self-evaluation to an external, expert-built benchmark.
The publication was paired with a Center for AI Safety / AI Frontiers essay titled "AIs Are Disseminating Expert-Level Virology Skills" (Dan Hendrycks & Laura Hiscott, 22 Apr 2025), which argues that the bottleneck for misuse is no longer information scarcity but information access[3].
Then the policy layer piled on. A Center for Strategic and International Studies (CSIS) analysis — Opportunities to Strengthen U.S. Biosecurity from AI-Enabled Bioterrorism: What Policymakers Should Know, by Georgia Adamson and Gregory C. Allen of the Wadhwani AI Center — frames the VCT as evidence a "threshold already crossed"[5]. The same report notes that the FY2026 U.S. budget proposal would cut NIST funding by ~$325 million (~30%) and eliminate roughly 500 NIST staff, while the administration's own AI Action Plan assigns NIST and its subordinate Center for AI Standards and Innovation (CAISI) the lead role on biosecurity evaluations of frontier models[6]. CSIS calls this out as a self-defeating mismatch: the agency assigned to do the biosecurity-AI work is being defunded while the capability is rising.
Three reasons this story is more than an LLM-benchmark headline.#
1. Pandemic preparedness is, among other things, a knowledge-distribution problem. The classic worry is a novel pathogen with pandemic potential. The less classic, increasingly urgent worry is a non-expert with high intent and a laptop. If the practical, troubleshooting side of virology is now something a frontier model can do at the 94th percentile of human experts, the marginal attacker no longer needs a PhD to clear the procedural hurdles of many wet-lab steps. That's exactly the threshold the Forecasting Research Institute's expert survey used when it tied a 1.5% annual risk of a 100,000+-death human-caused epidemic to "LLMs matching the performance of a top-performing team of virologists on a virology troubleshooting test" — a threshold the VCT shows has already been crossed[7].
2. The defender side gets the same windfall — and has to move faster. The same model that can troubleshoot an opaque plaque assay can also help a public-health lab triage a metagenomic dark-matter hit, a surveillance team interpret a recombination signal, or a vaccine group redesign an immunogen against a drifted variant. The week the VCT was making news, the WHO was simultaneously dealing with a Bundibugyo Ebola outbreak in DRC & Uganda (declared a Public Health Emergency of International Concern (PHEIC), with 2,011 confirmed cases and 754 deaths in DRC as of mid-July 2026), a Hantavirus outbreak linked to cruise-ship travel, and a Cyclospora outbreak in the U.S. tied to shredded iceberg lettuce from Mexico [8][9][10]. Protein language models like CoVFit, LucaVirus, and ESM-2-based tools such as VirHostPRED are already being used to predict variant fitness, host range, and antibody escape[12][13][14]. The question is no longer whether AI helps the good guys — it demonstrably does — but whether the good guys can absorb those tools faster than the threat surface expands.
3. Governance is the bottleneck, not capability. CSIS's three recommendations are deliberately modest: fund NIST and CAISI to do the biosecurity-AI work the AI Action Plan already assigned them; have CAISI lead formal evaluations of frontier biological design tools (BDTs) with support from the international AI Safety Institute network; and direct OSTP to build an AI-enabled DNA-synthesis screening system that can catch novel, AI-generated sequences that current static lists miss[5]. None of this is moonshot. The bipartisan interesting part is that the same week a benchmark showed LLMs clearing the expert line, the agencies assigned to evaluate that risk were facing proposed budget cuts and a leadership vacancy[6]. The pandemic-preparedness community should read that as a coordination problem, not a science problem.
The limitations#
This is a benchmark, not a wet-lab outcome study. VCT is a 322-question multiple-response test. It measures whether a model can pick the right troubleshooting step, not whether a human equipped with the model can execute a real protocol to a real outcome. The Biosecurity Handbook summary of pre-registered randomised evidence (Hong et al., 2026, preprint) found that mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures versus internet access (5.2% vs 6.6%, P=0.759), although they helped with intermediate steps[15]. So the VCT tells you the model knows the answer; the RCT tells you the human doing the procedure still mostly doesn't. The gap between those two is the actual attack surface.
The VCT itself is labelled as a preprint on arXiv (arXiv:2504.16137v2, 29 Apr 2025) — peer review is in progress / mixed across venues. Treat the percentile numbers as the authors' best characterization, not a refereed consensus[1].
The benchmark is asymmetrically hard for the humans. VCT experts took a "tailored question-set" in their own sub-field with internet access, and still averaged 22.1%. The models took the full benchmark. Apples-to-oranges concerns have been raised in commentary; the SecureBio authors argue the comparison remains informative because the models still beat the matched-expert human baseline. Readers should hold that caveat in mind.
"Dual-use" cuts both ways. The same authors and the CSIS report are explicit that this capability is also defensively useful: pandemic surveillance, antibody design, vaccine antigen redesign, genomic-dark-matter triage. The policy debate is not "should this exist" but "how do we get the upside without the catastrophic downside."
No medical or scientific advice. The model answers in the benchmark are research and educational signals. They are not validated laboratory procedures. None of this post is a substitute for trained virologists, BSL-rated facilities, or institutional biosafety committees.
Ethics, briefly. Authors at SecureBio and CAIS have called for "thoughtful access controls" on frontier models — the same kind of discussion now active at the U.S. Center for AI Standards and Innovation, the U.K. AI Safety Institute, the EU AI Act's general-purpose-AI obligations, and the WHO Hub for Pandemic and Epidemic Intelligence[16]. The dual-use character of these results means the responsible move is to publish the benchmark (so defenders can measure) and gate the most dangerous capabilities behind access controls (so attackers can't trivially consume them).
FAQs#
Q1. Did AI actually "beat" virologists? On a 322-question troubleshooting benchmark, OpenAI's o3 scored 43.8% versus an expert baseline of 22.1% on questions in each expert's own sub-field — placing o3 in the 94th percentile. That's a meaningful capability signal, not a license to do benchwork unsupervised[1].
Q2. Is this peer-reviewed? The VCT paper is a preprint on arXiv (arXiv:2504.16137v2). The OpenAI o1 System Card is a corporate technical report[4]. The CSIS report is a policy paper from the Wadhwani AI Center[5]. Treat each as the type of source it is.
Q3. Does this mean LLMs can now help someone build a pandemic pathogen? The peer-reviewed-benchmark signal says the information barrier to many troubleshooting steps is lower than it used to be. The pre-registered RCT evidence (Hong et al., 2026, preprint) says novices with an LLM still don't complete complex lab procedures significantly better than novices with internet access[15]. The two together suggest the operational barrier is real but eroding; whether it has fallen far enough to enable an attack is debated in the CSIS report, which settles on "moderate confidence" that commercial AI is "on the cusp" of meaningful novice uplift[5].
Q4. What's a "biological design tool" (BDT)? A narrower category of AI than LLMs — purpose-built systems for molecular engineering (protein/nucleic-acid design, codon optimization, etc.). CSIS flags BDTs as the second, longer-tail risk in their threat model; capability is harder to assess publicly[5].
Q5. What should a working virologist, public-health official, or student take away? Three things: (1) Frontier LLMs are now a legitimate, free troubleshooting adjunct for routine experimental problems — use them, with the usual lab-safety and data-confidentiality caveats. (2) The same capability argues for tighter institutional access controls, not looser ones, on the most capable models. (3) Pandemic preparedness is increasingly a governance and evaluation problem, not a sequencing-and-PCR problem — the bottleneck has moved.
Q6. Is there a medical advice angle here? No. This post is about research capability and pandemic preparedness. It is not medical advice, not a substitute for clinicians, and not a recommendation to use any AI system for self-diagnosis or self-treatment.
Primary-source links#
[1]: Götting, J., Medeiros, P., Sanders, J. G., Li, N., Phan, L., Elabd, K., Justen, L., Hendrycks, D., & Donoughe, S. (2025). Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark. arXiv preprint arXiv:2504.16137v2 (29 Apr 2025). https://arxiv.org/abs/2504.16137
[2]: SecureBio. AIs can provide expert-level virology assistance. (Substack explainer accompanying the VCT paper.) https://securebio.substack.com/p/ais-can-provide-expert-level-virology
[3]: Hendrycks, D., & Hiscott, L. (2025, 22 April). AIs Are Disseminating Expert-Level Virology Skills. AI Frontiers / Center for AI Safety. https://www.ai-frontiers.org/articles/ais-are-disseminating-expert-level-virology-skills
[4]: OpenAI. (2024). OpenAI o1 System Card. https://openai.com/index/openai-o1-system-card/
[5]: Adamson, G., & Allen, G. C. (2025, August). Opportunities to Strengthen U.S. Biosecurity from AI-Enabled Bioterrorism: What Policymakers Should Know. CSIS Wadhwani AI Center. https://www.csis.org/analysis/opportunities-strengthen-us-biosecurity-ai-enabled-bioterrorism-what-policymakers-should
[6]: Emerging Tech Lab Newswire. (2025). AI Models Breach Expert Virology Benchmarks as U.S. Biosecurity Oversight Thins (summary of CSIS report). https://emerging-tech-lab.com/press/ai-models-breach-expert-virology-benchmarks-as-u-s-biosecurity-oversight-thins
[7]: Forecasting Research Institute. Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards. https://forecastingresearch.org/research/llm-enabled-biorisk
[8]: World Health Organization. (2026, 3 July). Ebola disease caused by Bundibugyo virus, Democratic Republic of the Congo & Uganda — Disease Outbreak News. https://www.who.int/emergencies/disease-outbreak-news/item/2026-DON613
[9]: Hantavirus outbreak linked to cruise ship travel, Multi-locations https://www.who.int/emergencies/disease-outbreak-news/item/2026-DON611
[10]: U.S. Centers for Disease Control and Prevention. (2026, July). Investigation Update: Cyclospora Outbreak, July 2026. https://www.cdc.gov/cyclosporiasis/outbreaks/07-26/investigation.html
[11]: UK GOV.UK. (2026). Outbreaks under monitoring: week 28 https://www.gov.uk/government/publications/outbreaks-under-monitoring-in-2026/outbreaks-under-monitoring-week-28-week-ending-12-july-2026
[12]: Ito, J., et al. A protein language model for exploring viral fitness landscapes (CoVFit). PMC. https://pmc.ncbi.nlm.nih.gov/articles/PMC12075601/
[13]: LucaVirus: a unified nucleotide–protein language model for viral sequence analysis. National Science Review, 13(14), nwag376 (17 June 2026). https://academic.oup.com/nsr/article/13/14/nwag376/8709798
[14]: Protein language models enable accurate viral host range prediction (VirHostPRED). Sci Rep 16, 7606 (25 Feb 2026). https://pubmed.ncbi.nlm.nih.gov/41741511/
[15]: LLMs and Information Hazards. The Biosecurity Handbook. https://biosecurityhandbook.com/ai-biosecurity/llms-info-hazards.html (cites Hong et al., 2026, preprint)
[16]: WHO Hub for Pandemic and Epidemic Intelligence. (2026, 17 July). July 2026 update from the WHO Hub for Pandemic and Epidemic Intelligence. https://pandemichub.who.int/news-room/news/17-07-2026-july-2026-update-from-the-who-hub-for-pandemic-and-epidemic-intelligence
[17]: ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity. arXiv preprint arXiv:2606.11150. https://ar5iv.labs.arxiv.org/html/2606.11150