Biology
The 50-Nucleotide Question: Can DNA Synthesis Screening Survive AI Protein Design?
A new peer-reviewed study cut AI-redesigned toxin genes into fragments and fed them to four commercial DNA screening tools. Two caught fragments as short as 50 nucleotides. Here's what that means ten weeks before a US screening deadline — and why the authors are still worried.
If you want to build a dangerous protein, you don't need a lab full of equipment. You need a DNA sequence and a credit card. Commercial synthesis providers will print the gene and ship it to you.
The thing standing between that order and a biohazard is a piece of software called biosecurity screening software, or BSS. It compares your order against a list of sequences of concern — genes encoding toxins, viral proteins, the ingredients of things nobody should be assembling in a garage — and flags the matches.
That defence has a known weakness, and last October a team at Microsoft published it in Science: AI protein design tools can rewrite a toxin's amino acid sequence until it barely resembles the original, while keeping the three-dimensional fold that gives the protein its function (Wittmann et al., Science, 2 Oct 2025). The screeners were looking for sequence similarity. The AI had removed the sequence similarity and left the biology intact.
The same team spent the following year patching the tools with four commercial DNA synthesis companies. On 13 July 2026, they published the stress test of those patches — and the results are better than you'd expect, in a way that should still make you uncomfortable.
What happened#
The new paper, The limits of sequence-based biosecurity screening tools in the age of AI-assisted protein design, appeared in Frontiers in Bioengineering and Biotechnology on 13 July 2026. It's peer-reviewed, not a preprint, and the author list is unusual: Microsoft's chief scientific officer's team alongside people from Twist Bioscience, Integrated DNA Technologies, Aclid, Battelle, RTX BBN, the University of Birmingham, and IBBIS. That's the defenders and the vendors publishing together, which is not how biosecurity papers usually work.
The attack they modelled is straightforward. Suppose a screening tool now catches your AI-redesigned toxin gene when you order the whole thing. What if you don't order the whole thing? What if you order it in pieces, from different vendors, and assemble it yourself?
So they took AI-redesigned versions of proteins of concern that the patched tools had successfully flagged at full length, chopped them into fragments, and ran the fragments through four screening tools. They started at 200 nucleotides — the window size the 2024 US Framework for Nucleic Acid Synthesis Screening specifies — and worked down in 25-nucleotide steps to 25 nucleotides, which is seven or eight amino acids.
Two of the four tools held up at 50 nucleotides with no further modification, matching the performance they showed on full-length sequences. The other two had problems in opposite directions: one struggled to catch true positives as fragments got shorter, while the other flagged too many harmless sequences. Newer versions of both tools performed better.
The paper's own summary: the tested tools exceeded what the US framework currently asks of them.
There's a second result buried in the figures that matters more than the headline. The researchers plotted flagging rates against two measures of how much a redesigned fragment still resembles the original: percent identity, and longest common subsequence (LCS), the longest stretch of shared amino acids. Every tool has a cliff. One provider's tool flagged nearly 100% of fragments sharing a subsequence longer than 20 amino acids, then fell toward zero once identity dropped past roughly 30%. Another only failed when LCS fell below about 10 and identity fell below 50%.
Different cliffs, but all of them cliffs. Sequence-based screening works until the sequence stops being similar enough, and then it stops working. The question isn't whether the defence has a limit. It's how far today's design tools are from reaching it.
Why it matters#
The defence is holding, for now, and we know roughly why. A companion wet-lab study — a bioRxiv preprint, not yet peer-reviewed — generated AI-redesigned variants of three proteins chosen as escalating difficulty tiers (PDZ3, URA3, and T7 RNA polymerase) using ProteinMPNN, EvoDiff-MSA, and EvoDiff-Seq, then carried 300 of them — 100 per target — into the lab to test whether they actually worked. The finding: the design tools available in early 2024 couldn't reliably rewrite a protein enough to evade screening while keeping it functional. There's a tradeoff between divergence and function, and current models sit on the wrong side of it for an attacker. The July paper leans on this, noting that today's design tools "appear to be largely incapable of producing the extremely divergent (<30% sequence identity) sequences" that would break the screeners.
That's a capability claim about 2024-era models. Capability claims expire.
The replacement technology doesn't exist yet. The paper is blunt about the path forward. For longer sequences, you can fold a design in silico and compare its predicted structure to known threats, which is a better proxy for function than sequence matching. For a 50-nucleotide fragment in isolation, you can't. You'd need to predict function directly from a short, heavily mutated stretch of sequence, and the authors write that while early work on function-prediction and context-aware screening has begun, "none are currently operationally deployable." Their suggestion is that reliable calls on short fragments will require broader context — other sequences in the same order, other orders, information about the customer — rather than sequence analysis alone.
Screening coverage is thin in a specific, measurable way. Epoch AI, commissioned by Sentinel Bio, catalogued 1,196 biological AI models released since September 2024 in a report published 20 February 2026. Only 3.2% carry any documented safeguards. Only 2.5% have a documented pre-release risk assessment. The 35% figure for "notable" models is misleading on its own, because it's carried almost entirely by frontier LLMs, 95% of which have safeguards — against 1.4% of non-LLM biological AI models. Protein engineering and small molecule design together account for about half the database, and 253 of the 1,196 models are fine-tunes of existing models, most commonly the ESM-2 protein language model family.
So the ecosystem's structure is: a small number of heavily-guarded general models, and a long tail of specialised design tools built on shared open foundations with essentially no documented safety review. DNA synthesis screening is the chokepoint precisely because the model layer isn't one.
The policy layer is mid-rewrite. The 2024 White House Office of Science and Technology Policy (OSTP) Framework, released 29 April 2024, tied compliance to federal research funding: take US life sciences money, buy only from screened providers. Per the July paper, it asks tools to screen 200-nucleotide windows now and to handle 50-nucleotide fragments by October 2026 — roughly ten weeks from today. Meanwhile, the Administration for Strategic Preparedness and Response (ASPR)'s own page carries a notice that federal agencies will revise or replace that framework under the May 2025 Executive Order on Improving the Safety and Security of Biological Research, with no replacement published yet. In the Senate, S.3741, the Biosecurity Modernization and Innovation Act of 2026, introduced 29 January 2026 by Cotton and Klobuchar, would move screening from a funding condition to a Commerce Department regulation. It sits in committee.
A deadline is approaching for a framework that is being replaced by something not yet written, while the evidence base for what the deadline should say is a paper published three weeks ago.
The capability side isn't slowing. In May 2026, Biohub released an open protein "world model" and reported to Axios that it computationally designed binders against five cancer and immunology targets in days rather than years, with hit rates of 36–88% for compact minibinders. Nothing about that work is threatening. It's the same capability curve, pointed at medicine. Design tools that can reliably produce functional, sequence-divergent proteins are exactly what the field is trying to build, for good reasons, and they are exactly what breaks sequence-based screening.
The limitations#
I want to be precise about what this study does and doesn't show.
It's computational. The July fragment study, like the original Science work, was done on computers. No proteins were synthesised or tested for toxicity. Whether the redesigned sequences would actually function as toxins is a separate question, addressed partly by the wet-lab preprint, which used deliberately benign proxy proteins rather than real threat agents.
Four tools, anonymised. The paper reports results as "Provider N, Tool X" without naming which commercial screener is which. Reasonable for security, but it means a synthesis customer can't act on the findings. The paper notes the spread at 50 nucleotides was wide, roughly 0.5 to 0.8 in the Matthews correlation coefficient, and flags a real problem: a tool tuned for conventional threats may be worse against engineered ones, and providers have no clear basis for choosing between those risk profiles.
The threat model is one attacker strategy. Fragmenting an order is a plausible evasion route. It isn't the only one, and the study doesn't claim to cover the space.
Screening is one layer. Order screening sits alongside customer verification, know-your-customer requirements, export controls, and the practical difficulty of doing anything with a synthesised gene once you have it. A gap in one layer is not a pathway to a pandemic. This post is about the health of one control, not a risk assessment of anything.
FAQs#
Should I be worried about this? Not about your health, today. This is a story about infrastructure resilience, not an outbreak. The finding is that a safeguard is currently working and has a known expiry condition.
Is this "AI can design bioweapons"? No. The studies test whether AI-modified sequences slip past a screening filter. That's a question about the filter. Whether the resulting proteins would work is largely unresolved, and the one experimental study that looked at it — a preprint, using safe proxy proteins — found that current design tools can't reliably keep function while achieving evasion.
Why not just make screening stricter? False positives. One of the four tools in this study had its score dragged down by flagging benign sequences. Every false positive is a legitimate researcher's order held up. Screening has to work at commercial scale on mostly harmless traffic.
Who decides what counts as a sequence of concern? In the US, the framework is built on the 2023 Health and Human Services (HHS) guidance, and it applies as a condition of federal research funding. That structure is under revision. The International Gene Synthesis Consortium, whose members co-authored this paper, maintains industry screening practices independently.
What would actually fix this? Screening that predicts what a sequence does rather than what it resembles, combined with order- and customer-level context. The authors say this doesn't exist in deployable form yet. That's the honest answer, and it's why the paper closes by urging that AI design capabilities be continually re-benchmarked against screening performance rather than assessed once.
Sources#
- Wittmann BJ, Wheeler NE, Murphy ST, Mitchell T, Rife Magalis B, Gemler BT, Flyangolts K, Diggans J, Clore A, Beal J, Bartling C, Alexanian T, et al., Horvitz E. The limits of sequence-based biosecurity screening tools in the age of AI-assisted protein design. Frontiers in Bioengineering and Biotechnology, 13 July 2026. doi:10.3389/fbioe.2026.1858951 (peer-reviewed)
- Wittmann BJ, et al. Strengthening nucleic acid biosecurity screening against generative protein design tools. Science 390(6768):82–87, 2 October 2025. (peer-reviewed) — PubMed record
- Ikonomova S, et al. Experimental evaluation of AI-driven protein design risks using safe biological proxies. bioRxiv, 2025. Preprint — not peer-reviewed.
- Atanasov D, Zanichelli N, Denain J-S. Expanding our analysis of biological AI models. Epoch AI, 20 February 2026 (commissioned by Sentinel Bio).
- ASPR / HHS. 2024 OSTP Framework for Nucleic Acid Synthesis Screening — includes revision notice under the May 2025 Executive Order.
- Executive Order: Improving the Safety and Security of Biological Research, 5 May 2025.
- S.3741 — Biosecurity Modernization and Innovation Act of 2026, 119th Congress, introduced 29 January 2026.
- Science News. AI-designed proteins test biosecurity safeguards.
- Axios. Zuckerberg's Biohub unveils AI "world model" of proteins, 27 May 2026.