A computer drafted these antibodies before anyone pipetted a drop#
For most of the history of antibody medicine, finding a good antibody has meant finding it in a living system. You immunise an animal, or you screen a vast library of antibody variants, and you hope something sticks to your target. Now several teams say they can start from a blank page. They describe the target, ask a model to draft binders, and test a few dozen candidates in the lab. In one preprint, a team reported that its model produced working binders for half of 52 targets in a single round of testing, with no known binders for those targets in the Protein Data Bank (Chai Discovery, 2025, not peer reviewed).
That is a striking claim, and it deserves some care. A preprint is a manuscript that has not yet been through peer review, so it is a claim rather than a settled result. The peer-reviewed record is more modest, and in some ways more interesting. In 2025, David Baker's group published in Nature the first atomically accurate design of antibody fragments aimed at chosen sites on a target, but its success rates were low and often needed thousands of designs to be screened (Bennett et al., Nature, 2025). In June 2026, a Stanford-led team published a method called Germinal in Nature Biotechnology that needed only 43 to 101 designs per target to find working antibodies (Mille-Fragoso et al., Nature Biotechnology, 2026).
So the field has moved quickly from "can it be done at all?" to "how reliably, how cheaply and how safely?". This article walks through how a protein language model fits into that story, what the evidence supports, and where the hype runs ahead of the data.
Background: what a protein language model is, and why antibodies are a hard target#
Proteins as text#
Proteins are chains of amino acids, and each amino acid can be written as a letter. A protein language model is trained like the chatbots many of us use, except that the text is protein sequences. The model learns which letters tend to appear together and in which contexts, and in doing so it absorbs patterns that evolution has already tested. A 2023 paper in Science showed that scaling such models up to 15 billion parameters let them predict atomic-level protein structure directly from a sequence, and the authors used this to predict structures for more than 617 million metagenomic proteins (Lin et al., Science, 2023). That family of models is usually called ESM, and "ESM protein language model" is now a common search term in its own right.
Language models can also write. ESM3, a later model that reasons over sequence, structure and function together, was prompted to generate fluorescent proteins, and one bright example sat at only 58% sequence identity to known fluorescent proteins (Hayes et al., Science, 2025). The authors estimated this was equivalent to simulating 500 million years of evolution (Hayes et al., Science, 2025). The point for our story is that a model trained on natural sequences can generate functional proteins that natural selection never got round to producing.
What an antibody actually is#
An antibody is a Y-shaped protein that the immune system makes to grab a specific target. The grabbing is done by loops at the tips called complementarity-determining regions, or CDRs. These loops are the hard part to design, because tiny changes in them can switch binding on or off. Antibody discovery also has to deal with developability problems such as poor solubility, clumping and immune reactions against the protein itself (Shuai et al., Cell Systems, 2023).
Two antibody formats come up repeatedly in the new papers. A VHH, or nanobody, is a small single-chain antibody fragment. An scFv, or single-chain variable fragment, is a larger fragment that joins the two binding halves of a conventional antibody into one chain. Both formats were tested in the Germinal work (Mille-Fragoso et al., Nature Biotechnology, 2026).
Two kinds of AI, working as a pair#
Modern antibody design usually combines two different tools. One is a structure predictor or structure generator, which asks whether a proposed antibody would fold into a shape that fits the target. The other is a language model, which asks whether the sequence looks like something a real antibody could be. IgLM, for example, was trained on 558 million antibody sequences and can redesign variable-length stretches of a sequence, including CDR loops, with better predicted developability (Shuai et al., Cell Systems, 2023). The language model keeps the design realistic, and the structure tool keeps it physical.
How the first atomically accurate designs were made#
The Baker lab's Nature paper fine-tuned its RFdiffusion model, a generative network that builds protein structures, so that it could design antibody fragments against user-chosen sites on a target. The designs were then screened by yeast display, a laboratory method in which candidate proteins are shown on the surface of yeast cells and the ones that stick to the target are picked out (Bennett et al., Nature, 2025).
The structural accuracy was real. Cryo-electron microscopy, a technique that images frozen molecules at near-atomic detail, showed that a designed nanobody bound influenza haemagglutinin at 3.0 angstroms resolution with a CDR3 loop deviation of 0.8 angstroms from the design (Bennett et al., Nature, 2025). For a loop that is notoriously floppy, that is a remarkable match between prediction and reality.
The binding strength was more ordinary. The best influenza nanobody bound at 78 nM, and other designs ranged from hundreds of nanomolar to micromolar, which is useful as a starting point but weaker than many medicines (A lower number means tighter binding: nanomolar is good, picomolar is excellent). Success rates for nanobodies were between 0 and 2% across the targets tested, and the authors screened more than 9,000 designs per target. Affinity maturation, a lab process that mimics evolution to improve a binder, then raised affinity roughly 100-fold in some cases (Bennett et al., Nature, 2025).
The structures were right, but the pipeline still leaned heavily on screening to find the winners. The model could draw a convincing antibody. It could not yet tell you which of its drawings would work.
Where the protein language model enters: Germinal and the push for efficiency#
Germinal tackles the screening burden. It combines AlphaFold-Multimer, which predicts how two proteins fit together, with IgLM, which keeps the sequence antibody-like, and it refines designs in a loop until both tools are satisfied (Nature Methods research highlight, 2026). The CDRs are designed from scratch on a chosen framework, and the team validated both nanobody and scFv formats (Mille-Fragoso et al., Nature Biotechnology, 2026).
The reported numbers are better than the earlier Nature results. Across four targets (PD-L1, IL-3, IL-20 and BHRF1), success rates ran from 4 to 22% with 43 to 101 designs tested per target, and several antibodies reached sub-micromolar to nanomolar affinity (Nature Methods research highlight, 2026). The team also used cryo-EM and epitope mapping to check the designs, and it released its code openly (Mille-Fragoso et al., Nature Biotechnology, 2026).
The authors are upfront about a limit: testing covered only four targets, so how well the approach generalises is still open (Mille-Fragoso et al., Nature Biotechnology, 2026). That caveat matters more than the headline percentage. A method that works on four targets may or may not work on the fifth.
The commercial and preprint wave, read with care#
Around the peer-reviewed papers sits a fast-moving layer of preprints and company reports. They are worth reading because they show where the field is heading, and they are worth doubting because none has yet had independent expert scrutiny.
Chai Discovery's Chai-2 preprint reported a 16% hit rate across its tests, described as more than 100 times better than earlier computational methods, using 20 or fewer designs per target (Chai Discovery, 2025, not peer reviewed). A separate preprint, mBER, designed more than 1.15 million nanobody sequences against 436 human cell-surface targets, tested 145 targets in phage display (a screening method similar in spirit to yeast display), and found elevated hit rates for 65 of them, with rates as high as 38% for the best epitopes under strict filtering (mBER, 2025, not peer reviewed). Its authors note that success depends heavily on which epitope is chosen (mBER, 2025, not peer reviewed). A third preprint reported antibodies against the receptors CXCR4 and CXCR7, including top designs at 370 picomolar and some computationally designed agonists, after repeated rounds of design and filtering (Nabla Bio, 2025, not peer reviewed).
The chart below puts the headline numbers side by side. It should be read as a map of claims, not a league table, because each team measured success differently, used different targets and filtered differently.
Figure: highest hit rate reported in each source. Sources: Bennett et al., 2025; Chai Discovery, 2025 (not peer reviewed); Nature Methods, 2026; mBER, 2025 (not peer reviewed). The RFdiffusion bar shows the top of the 0 to 2% range for nanobodies.
One more caution about the chart. A "hit" usually means that a designed protein stuck to the target in a screening assay. It does not mean the antibody is potent, stable, safe or manufacturable. Those properties come later, and they decide whether a molecule becomes a medicine.
What could still go wrong: immunogenicity, developability and biosecurity#
The drug is only as good as the immune system's reaction to it#
Even approved antibody drugs can provoke the body to make antibodies against them. A 2026 review of rituximab, a widely used antibody therapy, found that anti-drug antibodies were linked in some patients to lower drug levels, earlier return of B cells, more relapses and hypersensitivity reactions, although not every case was clinically relevant (Allinovi et al., Frontiers in Immunology, 2026). A model that writes antibody-like sequences does not automatically write sequences the immune system will ignore. Language models like IgLM aim to improve predicted developability, but those predictions are made in silico, meaning on a computer, and still need to be tested in people (Shuai et al., Cell Systems, 2023).
Dual use is a design problem too#
The same tools that design useful proteins can, in principle, redesign dangerous ones. A 2025 Science study led by Microsoft researchers tested whether open-source AI protein design software could create variants of proteins of concern that evaded the screening used by DNA synthesis companies, and found that AI-redesigned sequences could not be detected reliably by the tools then in use (Wittmann et al., Science, 2025). The team then built and deployed patches that greatly improved detection of synthetic homologues likely to keep wild-type-like function. The study also argued that DNA synthesis is a choke point in AI-assisted protein engineering, which makes order screening a sensible place to focus protection. That paper concerns protein design in general rather than antibodies specifically, but it is the right frame for anyone building these tools.

What happens next#
Three practical questions will decide how far this goes.
The first is generality. Germinal covered four targets and Nature's RFdiffusion work covered a handful, and both groups say that broader testing is needed (Mille-Fragoso et al., Nature Biotechnology, 2026), (Bennett et al., Nature, 2025). The preprints cover many more targets, but they have not yet been independently scrutinised. If their results survive peer review, the picture will look different.
The second is potency and drug-likeness. Designs that bind at nanomolar strength are a good start, yet the Nature paper noted that some designed binders were too weak for a demanding application such as CAR-T cell killing, and that filtering by structure prediction had limited power (Bennett et al., Nature, 2025). Better scoring of designs, before the lab step, is probably where progress will come from.
The third is clinical evidence. There is no peer-reviewed report of a fully de novo, AI-designed antibody tested in patients, so the clinical question is open. Until that data exists, the fair reading is that AI has made antibody discovery faster to start and cheaper to explore, while the slower, riskier work of proving a drug safe and effective remains as hard as before.
For researchers and founders, the practical advice is plain. Treat open tools such as Germinal as a way to generate starting points, budget for experimental validation of every design, and read preprint hit rates as claims to be tested. For readers outside the field, the headline is that a protein language model can now help write the first draft of an antibody, and that the editing still happens in a lab.
Frequently asked questions#
What is a protein language model? It is an AI model trained on millions of protein sequences, in the way a chatbot is trained on text, so that it learns the patterns evolution has produced. Models in the ESM family can predict protein structure from sequence alone and can generate new, functional proteins (Lin et al., Science, 2023), (Hayes et al., Science, 2025).
What does "de novo antibody design" mean? It means designing an antibody on a computer without starting from an existing antibody against the target. The 2025 Nature paper described this as a capability that had not previously existed for epitope-specific antibodies (Bennett et al., Nature, 2025).
How high are the success rates? They vary widely. Peer-reviewed work reported 0 to 2% for nanobodies in one study and 4 to 22% for Germinal across four targets (Bennett et al., Nature, 2025), (Nature Methods research highlight, 2026). Preprints report higher figures, such as 16% for Chai-2 (not peer reviewed) (Chai Discovery, 2025).
Does a "hit" mean the antibody could become a drug? No. A hit means a design bound its target in a screening test. Drugs also need strong affinity, stability, low immunogenicity and reliable manufacturing, and designed binders have sometimes been too weak for demanding uses (Bennett et al., Nature, 2025).
Is any of this tool open to academics? Germinal's authors released open-source code and protocols, so groups can test it themselves (Mille-Fragoso et al., Nature Biotechnology, 2026).
Why do antibodies provoke immune reactions even when they are made by humans? Therapeutic antibodies can trigger anti-drug antibodies, which in some patients are linked to lower drug levels and more relapses, as seen with rituximab (Allinovi et al., Frontiers in Immunology, 2026).
Are there biosecurity risks? Yes, for AI protein design in general. A 2025 Science study found that AI-redesigned versions of hazardous proteins could slip past DNA synthesis screening, then showed that software patches greatly improved detection (Wittmann et al., Science, 2025).
Has an AI-designed antibody been tested in people? There is no peer-reviewed evidence of a fully de novo AI-designed antibody tested in patients, so the answer is not yet established in the literature reviewed here.
References#
- Mille-Fragoso LS, Driscoll CL, Wang JN, et al. Efficient generation of epitope-targeted antibodies with Germinal. Nature Biotechnology, 23 June 2026. DOI: 10.1038/s41587-026-03187-0
- Research highlight: AI-designed antibodies with Germinal. Nature Methods, 6 August 2026. DOI: 10.1038/s41592-026-03190-y
- Bennett NR, Watson JL, Ragotte RJ, et al. Atomically accurate de novo design of antibodies with RFdiffusion. Nature, 2025. DOI: 10.1038/s41586-025-09721-5
- Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 2023;379:1123-1130. DOI: 10.1126/science.ade2574 (PubMed: 36927031)
- Hayes T, Rao R, Akin H, et al. Simulating 500 million years of evolution with a language model. Science, 2025;387:850-858. DOI: 10.1126/science.ads0018 (PubMed: 39818825)
- Shuai RW, Ruffolo JA, Gray JJ. IgLM: Infilling language modeling for antibody sequence design. Cell Systems, 2023;14:979-989. DOI: 10.1016/j.cels.2023.10.001 (PubMed: 37909045)
- Wittmann BJ, Alexanian T, Bartling C, et al. Strengthening nucleic acid biosecurity screening against generative protein design tools. Science, 2025;390:82-87. DOI: 10.1126/science.adu8578 (PubMed: 41037625)
- Allinovi M, Lomi J, Accinno M, et al. Incidence and implications of anti-rituximab antibodies in nephropathies and other immune-mediated diseases: a narrative review. Frontiers in Immunology, 2026;17:1896560. DOI: 10.3389/fimmu.2026.1896560 (PubMed: 42812566)
- Chai Discovery Team. Zero-shot antibody design in a 24-well plate. bioRxiv, July 2025. (not peer reviewed). 10.1101/2025.07.05.663018
- mBER: Controllable de novo antibody design with million-scale experimental screening. bioRxiv, 2025. (not peer reviewed). 10.1101/2025.09.26.678877
- Nabla Bio. De novo design of antibodies against CXCR4, CXCR7 and other targets by scaling test-time compute. bioRxiv, 2025. (not peer reviewed). 10.1101/2025.05.28.656709