An AI wrote working viruses. So who is in charge?#

In September 2025, a team from Stanford University and the Arc Institute reported something that sounded like science fiction. They had asked genome language models, the same family of software that powers chatbots but trained on DNA, to write entire viral genomes from scratch. They then built the DNA in the lab and dropped it into bacteria. Sixteen of the designs came alive and infected their hosts (Arc Institute and Stanford)

These were bacteriophages, viruses that attack bacteria and cannot infect people. The researchers deliberately trained their models to leave out viruses that infect humans and other complex organisms (Science). That was a careful, responsible choice. It was also a choice made by the researchers themselves, not by any law.

Nobody is claiming that a rogue chatbot is about to cook up a pandemic. The real question is quieter and more practical: if software can write the instructions for a virus, which rule book applies, and who enforces it? The answer, as of October 2026, is a patchwork. Some pieces are strong. Several are voluntary. A few are missing entirely.

This post walks through the science that triggered the debate, the safety net that currently exists, and the gaps that policy researchers say remain. It is written for readers who care about the technology and want a clear view of the governance, not a scare story.

Background: what does it mean to "design" a virus with AI?#

A genome is the full set of genetic instructions in an organism, written in a four-letter chemical alphabet (A, C, G and T). A genome language model works much as a text-predicting chatbot does. It reads huge quantities of DNA and learns which "letters" tend to follow which. Ask it to continue a sequence and it will write plausible new DNA.

In the Stanford and Arc work, the models, called Evo 1 and Evo 2, were pointed at a small, well-studied phage called ΦX174, which has about 5,400 DNA letters and 11 genes. Researchers synthesised and tested around 300 AI-generated designs. Of 302 candidates, 285 were successfully assembled, and 16 turned out to be functional phages. Some of the designs were as little as 93% identical to natural sequences, and several outperformed the original virus in lysis speed, meaning how quickly they burst bacterial cells. A cocktail of the designs even overcame bacterial strains that had become resistant to the natural phage.

What the work does show is a direction of travel. Designing at the scale of a whole genome used to be out of reach. Now it has been done, once, in a controlled setting. Governance has to be built for where the technology is heading, not just where it stands today.

The safety net that exists today, and the hole in it#

Most of the world's biosecurity protection against synthetic DNA rests on one idea: check the order before you make the molecule. Companies that sell custom DNA compare each requested sequence against databases of known dangerous pathogens and toxins, and they check who is placing the order. This is called gene synthesis screening, and it is the main technical checkpoint between a digital design and a physical organism.

The industry body behind much of this is the International Gene Synthesis Consortium. It covers roughly 80% of global synthesis capacity, which leaves about a fifth of the market outside its voluntary system. Existing screening tools reportedly identify known sequences with around 95% accuracy (Arms Control Association).

The weakness is built into the method. Screening works by matching what it has seen before. It struggles with sequences that have never been seen, which is exactly what a generative model produces. Policy researchers writing in Frontiers in Bioengineering and Biotechnology in July 2026 make the same point: similarity-based screening weakens against codon optimisation and AI-generated variants (Kato et al. 2026). Codon optimisation, in plain terms, means rewriting DNA using different "spellings" that still produce the same protein.

The red-team test that rattled the screeners#

If the hole is theoretical, someone ought to test it. A team led by Microsoft's Eric Horvitz did exactly that, in a study published in Science. They used AI protein design tools to generate more than 75,000 variants of hazardous proteins and tested whether commercial screening software would catch them (NPR; Science).

The screening tools were not consistently able to catch the redesigned sequences. A software patch was developed and deployed, which improved detection, but it still missed a small fraction of the variants (NPR). The project involved researchers from Microsoft, the biosecurity company RTX BBN, and the DNA suppliers Twist Bioscience and IDT, among others (Microsoft Research).

The team did not run laboratory tests to confirm that the redesigned proteins would actually function, in part because of the legal and ethical sensitivities involved. And they restricted access to the underlying data rather than publishing it openly, which was described as a first for managing biosecurity risk in a scientific publication. It was a rare case of researchers, vendors and a journal agreeing on a disclosure procedure before the vulnerability became public, rather than after.

Who regulates what? A jurisdiction-by-jurisdiction map#

The "policy gap" in the title becomes concrete here. There is no single authority that oversees an AI model, the DNA it designs and the lab that builds it. Each jurisdiction covers a slice. The table below summarises the picture as described by recent policy analyses and official documents.

Jurisdiction or bodyWhat it does todayMain gap
United StatesFederal guidance encourages sequence screening and customer verification; the 2024 screening framework applied mandatory expectations to federally funded research (Frontiers, 2026; Arms Control Association)Not binding for privately funded work; the revised framework ordered in 2025 has not yet appeared (Congressional Research Service)
European UnionDual-use export controls and biosafety rules; AI Act obligations for the largest general-purpose models, including assessing risks such as help with chemical or biological weapons (Frontiers, 2026; Wilson Sonsini)No unified mandatory sequence-level screening rule (Frontiers, 2026)
United KingdomVoluntary screening guidance; AI Security Institute evaluating frontier model capabilities in biology (UK Government)Government is still weighing whether to put screening on a statutory footing (UK Government)
Biological Weapons ConventionInternational prohibition on biological weapons (Frontiers, 2026)No synthesis-specific screening mechanism or detailed operational rules (Frontiers, 2026)
Industry (IGSC)Voluntary screening protocols and customer checks (Congressional Research Service)Covers about 80% of global capacity (Arms Control Association)

Read down the right-hand column and a pattern appears. Almost every gap is the same one in a different outfit: the rules either apply to the wrong stage of the process, or they are voluntary, or they have not been finished.

who regulates an ai that designs a virus 2

Regulating the AI itself, not just the DNA#

So far we have looked at the physical end of the pipeline. A second approach is to regulate the AI models. The European Union's AI Act is the furthest along. Its code of practice for general-purpose AI, published on 10 July 2025, asks providers of the most capable models to identify and analyse systemic risks, with examples that include facilitating the development of chemical or biological weapons (Wilson Sonsini). Providers must assess and monitor those risks, apply mitigations such as filtering and monitoring, and report serious incidents. The obligations took effect on 2 August 2025, and a one-year grace period for code signatories ended on 2 August 2026, when fuller enforcement and fines became possible.

There is a catch. These rules are aimed at general-purpose models, the chatbot-style systems that millions of people use. A specialised biological design tool, trained on DNA or proteins, may not fit neatly into that category, even though it is arguably closer to the risk. That is an interpretation problem, and policy analysts have flagged design-stage tools as an area where oversight has not caught up (Frontiers, 2026).

In the United States, the picture runs through executive action and voluntary commitments. Executive Order 14292, signed in May 2025, directed the White House science office to revise policies on dual-use research and on nucleic acid synthesis screening (Congressional Research Service). A July 2026 policy on high-risk life sciences research followed, but the Office of Science and Technology Policy has not yet updated the 2024 screening framework as the order directed. Meanwhile, several AI companies have published their own safeguards. OpenAI released a governance framework addressing biological, chemical and cyber risks in May 2026, and Anthropic proposed a policy framework in June 2026 that would require frontier developers to test models for catastrophic risks and publish findings. Self-regulation can move quickly, but it is only as strong as each company's incentives.

Benchtop machines, borders and who bears the blame#

Two further problems make the gap wider than it first looks.

The first is the benchtop synthesiser. These are desktop machines that print DNA on site, so an order never passes through a company's screening desk. Enzymatic synthesis, a newer method, can now produce stretches of up to 120 DNA letters, up from about 50 previously (Arms Control Association). Mandatory standards currently apply mainly to federally funded research, so non-federally funded providers, manufacturers and customers face no binding requirements. A 2017 study showed the stakes of weak checks when researchers made horsepox virus from legally purchased DNA fragments.

The second is accountability. If a model designs something harmful, a third party synthesises it and a fourth uses it, who is responsible? Researchers describe this as a liability gap, particularly across borders. The same analysis warns that mandatory requirements differ around the world and that voluntary standards only bind the providers who sign up. It also stresses that lower-resource settings need technical help, not just rules, if global baselines are to work (Frontiers, 2026).

What are experts proposing?#

The Frontiers authors suggest assessing risk along four dimensions: design capability, access to synthesis, strength of governance, and clarity of liability. They pair these with four escalating tiers of response, from "controlled" to "high-risk". Their recommendations include keeping sequence screening but embedding it in wider governance, extending oversight upstream to design tools and downstream to benchtop devices, harmonising baselines across countries, and clarifying who is accountable at each step (Frontiers, 2026).

The United Kingdom has signalled it may follow part of this path. Its biological security strategy report says it will determine whether screening guidance should move onto a statutory footing, scope legislation to proportionately mitigate dual-use risks from AI-enabled biological capabilities, and publish a second AI Security Institute report on frontier capabilities in biology (UK Government).

Not everyone agrees on how tough rules should be. Heavy controls on biological design tools could slow useful work, such as phage therapies that fight antibiotic-resistant infections, which was a motivation for the Stanford and Arc project in the first place. Light-touch rules could leave a few unscreened routes open. The middle path in the sources above is proportionate oversight that scales with risk rather than one rule for everything.

What this means for researchers and founders#

If you build or use biological AI models, the practical takeaways are modest but real. Document what your training data excludes, as the Arc team did. Buy DNA from screened providers. Keep an eye on the EU's obligations if you serve European users, and on the evolving US and UK guidance. And if you find a weakness, consider a staged disclosure like the one used in the toxin study, rather than a full public dump.

Frequently asked questions#

Can AI really design a working virus? In a controlled experiment, yes, for a small bacteriophage. A Stanford and Arc Institute team reported 16 functional AI-designed phage genomes, a virus that infects only bacteria.

Is it legal to use AI to design a virus? It depends on the jurisdiction, the virus and the use. No single law bans designing viral sequences with AI, and the Biological Weapons Convention prohibits weapons but contains no synthesis-specific screening rules (Frontiers, 2026). This is general information, not legal advice.

What is gene synthesis screening? It is a check that DNA suppliers run on orders, comparing requested sequences against databases of dangerous pathogens and toxins and verifying customers (Congressional Research Service).

Why can AI-designed sequences slip past screening? Screening mostly matches against known sequences, and generative tools can produce new ones, as the Microsoft-led study showed with redesigned toxins (NPR).

Does the EU AI Act cover biological AI? Partly. Its general-purpose AI code of practice asks providers of the most capable models to assess risks such as facilitation of chemical or biological weapons, but specialised design tools may sit awkwardly outside that category (Wilson Sonsini; Frontiers, 2026).

Is the United States regulating this? Through executive action and funding conditions rather than a statute. Executive Order 14292 asked for revised screening and dual-use policies, but the 2024 screening framework has not yet been updated (Congressional Research Service).

What is a benchtop DNA synthesiser and why does it matter? It is a desktop machine that makes DNA on site, which can bypass provider-level screening, and binding standards for these devices are limited (Arms Control Association).

Should researchers worry about publishing biological AI work? They should think about it. The toxin study restricted access to its data as a way to share findings safely, an approach worth considering for sensitive work (NPR).

References#

  1. Arc Institute and Stanford University researchers. Generative design of bacteriophages with genome language models. Science, 2026. https://www.science.org/doi/10.1126/science.aec2657
  2. Wittmann B. et al. Strengthening nucleic acid biosecurity screening against generative protein design tools. Science, 2025. https://www.science.org/doi/10.1126/science.adu8578
  3. NPR. AI-designed proteins and biosecurity. 2 October 2025 (science journalism, used for reporting on the Science study). https://www.npr.org/2025/10/02/nx-s1-5558145/ai-artificial-intelligence-dangerous-proteins-biosecurity
  4. Microsoft Research. Paraphrase Project. https://www.microsoft.com/en-us/research/project/paraphrase-project/
  5. Kato S. E. et al. Strengthening global biosecurity for synthetic nucleic acid technology: from sequence screening to risk-based governance in the AI era. Frontiers in Bioengineering and Biotechnology, 2026; 14:1820001. https://www.frontiersin.org/journals/bioengineering-and-biotechnology/articles/10.3389/fbioe.2026.1820001/full
  6. Arms Control Association. Regulatory Gaps in Benchtop Nucleic Acid Synthesis Create Biosecurity Vulnerabilities. 24 November 2025. https://www.armscontrol.org/blog/2025-11-24/regulatory-gaps-benchtop-nucleic-acid-synthesis-create-biosecurity-vulnerabilities
  7. Congressional Research Service. Artificial Intelligence and Biosecurity Issues. https://www.everycrsreport.com/reports/IF13269.html
  8. UK Government. UK Biological Security Strategy: Implementation Report, July 2025 to July 2026. https://www.gov.uk/government/publications/uk-biological-security-strategy-implementation-report-july-2025-july-2026/uk-biological-security-strategy-implementation-report-july-2025-july-2026
  9. Wilson Sonsini. EU Releases Final Code of Practice for General-Purpose AI Models. (law-firm summary of the European Commission's code of practice, used because the primary text was summarised rather than read directly). https://www.wsgr.com/en/insights/eu-releases-final-code-of-practice-for-general-purpose-ai-models.html