AI
Two Doors Into Biology: OpenAI Hands Frontier Models to 100,000 Academics While Anthropic Bolts Its Own Shut
OpenAI is giving 100,000 academic researchers free frontier model access with 75+ life science skills. Anthropic's newest flagship refuses to explain mitochondria. Both companies claim to be managing the same biosecurity risk — and the gap between their answers tells you how unsettled AI-for-biology governance really is.
Yesterday, OpenAI announced it will give 100,000 academic researchers free access to its most capable models, complete with more than 75 life science skills covering genetics, genomics, sequencing, single-cell analysis, protein modeling, and drug discovery.
Seven weeks earlier, Anthropic shipped a flagship model that, according to press reports, declines to explain what mitochondria do.
Same underlying technology. Same stated concern — that frontier models are becoming capable enough in biology to matter for biosecurity. Opposite conclusions about what to do next. If you work anywhere near virology or pandemic preparedness, that divergence is the story worth your attention this week, more than any individual benchmark number.
What happened#
On July 29, 2026, OpenAI launched ChatGPT for Academic Researchers: free frontier model access for 100,000 researchers at selected degree-granting institutions with high research activity, beginning with 10,000 this summer at places including the Institute for Advanced Study and École normale supérieure, expanding through 2027. Approved researchers get GPT-5.6 Sol Pro, expanded deep research, higher usage limits, larger context windows, and four collaborator invites. The program sits inside a commitment of more than $250 million through 2027 for external scientific research.
The life sciences tooling is the part that matters here. OpenAI says researchers can use more than 75 life science skills plus connectors reaching scientific literature, public genomic and clinical databases, computational notebooks, and data platforms. This is not a chatbot with a biology hobby. It is an agentic research harness pointed at sequence data.
OpenAI also published usage figures: roughly 1.3 million people per week already use ChatGPT for advanced science and mathematics, generating about 8.4 million messages, and researchers in the top 20% of AI usage within their field are nearly twice as likely to hand the model tasks estimated to take four hours or more — about 7% of their requests versus 3.5% for peers.
On capability, OpenAI's GPT-5.6 announcement reports GPT-5.6 Sol Pro solving 31.5% of GeneBench Pro, a long-horizon genomics and quantitative biology evaluation. Plain Sol scores 28.7%, against 12% for GPT-5.5, 16% for Claude Opus 4.8, and 3.1% for Gemini 3.1 Pro Preview. On LifeSciBench, Sol reaches 59.9%.
Buried in a footnote on that same page is the contrast: Claude Fable 5 is excluded from the GeneBench Pro comparison because it "does not answer advanced biology questions and refuses the majority of questions in this eval."
That is not an accident of training. Reporting on Fable 5, which Anthropic released on June 9 as the first public model in its Mythos class, describes safety classifiers that detect prompts touching biology, chemistry, cybersecurity, and model distillation, then either refuse or route the query to the older Claude Opus 4.8. Journalists testing it found refusals on mitochondria, cell membranes, prions, and how mRNA vaccines work. Anthropic's stated rationale, per the same reporting, is that models capable of real scientific work could be misused for high-risk biological research, so the company chose to be deliberately over-conservative.
Two labs, one risk model, two doors: one propped open for verified academics, one welded shut for everyone.
Why it matters for virology and preparedness#
The honest reason this is hard is that the capability in question is not obscure. It is the ordinary competence of a working virologist.
The Virology Capabilities Test — a preprint from SecureBio and the Center for AI Safety, 322 multimodal questions on troubleshooting wet-lab virology protocols — found that expert virologists with internet access averaged 22.1% on questions inside their own sub-specialties, while OpenAI's o3 reached 43.8%, outperforming 94% of those experts. That result is over a year old now and was measured on a model two generations behind Sol.
The newer worry is agentic. ABC-Bench, a June 2026 preprint later presented as an ICML 2026 poster, tested agents on writing code for liquid-handling robots, designing DNA fragments for in vitro assembly, and evading DNA synthesis screening. PhD biologists with at least two years of coding experience averaged 24%; the top model, Grok 3, hit 53%. In wet-lab validation, scripts written by o4-mini-high ran on an OpenTrons robot and assembled DNA with the expected sequences.
So the capability is real, it transfers to hardware, and it is dual-use in the least convenient way: the skill that lets a model help you rescue a failed rescue experiment is the skill that lets it help someone else. There is no version of "protein modeling and sequence analysis for 100,000 people" that carves this out cleanly.
OpenAI's answer is architectural rather than absolute. The company says GPT-5.6 is more capable in biology than its predecessors but does not cross the Critical threshold in that category, and that testing suggests it "can support legitimate research but does not provide the end-to-end capability needed to create, engineer, or synthesize a highly dangerous novel threat." Safeguards are layered — model-level training, real-time checks, a reasoning monitor that reviews conversation context rather than relying on classifier flags alone, continuous monitoring, account-level enforcement — and were tested with roughly 700,000 A100-equivalent GPU hours of black-box automated red teaming plus a standing bio bug bounty.
Above that sits a separate tier. In May 2026 OpenAI launched Rosalind Biodefense, giving vetted developers and government partners access to GPT-Rosalind, its life-sciences reasoning model. Launch partners include Fourth Eon (function-based DNA synthesis screening), SecureDNA, and SecureBio Detection; government and allied partners include Lawrence Livermore National Laboratory, Johns Hopkins Applied Physics Laboratory, and CEPI, which is applying it to its 100 Days Mission for epidemic vaccine development. That page also notes that ChatGPT agent, in July 2025, was the first model OpenAI treated as High Capability in biology under its Preparedness Framework.
What is emerging, then, is a three-tier structure: broad public access with heavy guardrails, verified academic access with lighter friction, and trusted biodefense access to the most capable life-sciences models. Whether that stratification actually tracks risk — or merely tracks institutional prestige and paperwork — is an open empirical question, and nobody outside the labs can currently check.
Meanwhile the thing the whole apparatus is meant to prepare for keeps ticking along at its own pace. The WHO's Avian Influenza Weekly Update #1054, dated July 24, 2026, continues routine H5N1 monitoring; WHO's risk assessment covering 13 June to 7 July 2026 reported a third human H5N1 case in Bangladesh, again from Sylhet Division, with no sustained human-to-human transmission reported and overall public health risk from influenza viruses at the human-animal interface assessed as low. The models are getting better at virology faster than the surveillance systems they are supposed to support are getting better at anything.
The limitations#
31.5% is not a scientist. GeneBench Pro's 129 problems were, per secondary coverage of the benchmark, estimated by reviewers to take a human expert 20 to 40 hours each. Solving under a third of them, with maximum reasoning and Pro mode enabled, is a real capability jump and nowhere near a colleague you would trust unsupervised. Read the number as a slope, not a level.
Benchmarks measure benchmarks. VCT and ABC-Bench are preprints, not peer-reviewed literature, and both use constructed tasks with graded answers. ABC-Bench's wet-lab validation is a genuine step past pure Q&A, but three successful DNA assemblies on a liquid handler is a proof of concept, not an operational threat model. Equally, "does not cross the Critical threshold" is OpenAI grading its own homework against its own framework.
No weights, no independent check. As Axios reported, the program gives roughly the equivalent of a $200-per-month Pro account for a year but does not provide model weights or training data. OpenAI and Anthropic argue that withholding weights prevents misuse; researchers counter that limited access hampers independent evaluation, reproducibility, and safety work. Both are right, which is the problem. The program expands what scientists can do with the model while leaving the model itself closed.
The chokepoint downstream is still soft. Model-level safeguards matter less if synthesis screening is optional. Reporting on the current US picture describes DNA synthesis screening as largely voluntary and notes that a bipartisan 2026 bill to mandate it does not squarely address AI-designed sequences engineered to evade existing detection — the exact capability ABC-Bench was built to measure. Fourth Eon's function-based screening work under Rosalind Biodefense is aimed at that gap, which is encouraging, and voluntary industry initiatives are not a substitute for a rule.
Access is a selection effect. Eligibility runs through recognized, high-research-activity, degree-granting institutions. Genomic surveillance capacity is thinnest exactly where those institutions are thinnest. A program that accelerates 100,000 researchers at well-resourced universities is good; it is not the same as strengthening global outbreak detection, and it may widen the gap it is sometimes described as closing.
FAQs#
Is OpenAI giving researchers a model that could help build a pathogen? OpenAI's position is no: GPT-5.6 is assessed as not crossing the Critical biological threshold, and safeguards plus monitoring apply to academic accounts as elsewhere. The unresolved issue is that expert-level protocol troubleshooting — demonstrated in the VCT preprint over a year ago — is genuinely useful to both defenders and attackers, and no safeguard cleanly separates the two.
Is Anthropic's approach safer? It is more restrictive, which is not the same thing. Refusing to explain mRNA vaccines to a graduate student does not reduce the capability of open-weight models, and OpenAI's own GPT-5.6 write-up makes the counterargument explicitly for cybersecurity: overblocking creates a security risk of its own by disarming defenders while attackers use other tools. Whether that logic holds equally in biology, where the attacker's task is harder and slower, is a real disagreement among biosecurity researchers rather than a settled question.
Does this help pandemic preparedness? Indirectly, and mostly through the separate Rosalind Biodefense track — CEPI's 100 Days Mission, LLNL countermeasure work, screening infrastructure. The academic program is a general research accelerator, not a preparedness intervention.
Can I apply? Applications opened July 29 for qualifying researchers at eligible institutions, requiring verification of institutional affiliation and a description of intended scientific use. Details are on OpenAI's announcement page.
What should I watch next? Three things: whether GPT-5.6-class biology capability shows up in independent, peer-reviewed evaluations rather than lab self-reports; whether DNA synthesis screening becomes mandatory and function-based; and whether any of this reaches surveillance systems in the regions where novel spillovers are most likely.
Nothing here is medical advice. For H5N1 or any outbreak guidance, consult WHO, your national public health agency, or a clinician.
Sources#
- OpenAI, Accelerating scientific discovery with ChatGPT for Academic Researchers, July 29, 2026
- OpenAI, GPT-5.6: Frontier intelligence that scales with your ambition, July 9, 2026
- OpenAI, Strengthening societal resilience with Rosalind Biodefense, May 29, 2026
- Ina Fried, Axios, Exclusive: OpenAI offers 100,000 academics free ChatGPT access, July 29, 2026
- Preprint: Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark, arXiv:2504.16137 (SecureBio / Center for AI Safety)
- Preprint / ICML 2026 poster: ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity, arXiv:2606.11150
- SecureBio, Virology Capabilities Test
- WHO, Avian Influenza Weekly Update #1054, July 24, 2026
- WHO, Avian influenza A(H5N1) virus
- Crypto Briefing, Anthropic's Claude Fable 5 won't answer basic biology questions, and that's by design — secondary reporting on Fable 5 refusal behavior
- Singularity Hub, AI Can Now Design and Run Thousands of Experiments Without Human Hands, June 4, 2026 — secondary reporting on US DNA synthesis screening policy
- Congressional Research Service, Artificial Intelligence and Biosecurity Issues