Biotechnology
AI Beat the Lab at Antibody Design, But Just Once
The first blinded, wet-lab-validated benchmark for AI antibody design just landed in Nature Biotechnology. Here is what the models got right, and wrong.
The first blinded contest for AI antibody design#
For years, computational biologists have made a familiar promise: feed the algorithms enough data, and they will design better antibodies. This month that promise was put on trial. It was blinded, refereed, and checked at the bench. The verdict, published in Nature Biotechnology, is more interesting than either the boosters or the sceptics would have guessed. Artificial intelligence won one round outright, and lost two others to random chance.
AIntibody results#
On 19 August 2026, Nature Biotechnology published the full results of AIntibody, described by its organisers as the first international, blinded benchmark for AI-designed antibodies. Twenty-nine organisations, spanning universities, non-profits, AI start-ups, biotechnology firms and large drug and technology companies, submitted 511 antibodies designed or ranked by their computational models. Each one was then built and tested in the same laboratory under identical conditions, and no team saw a single experimental result until the competition had closed.
The models were set three tasks, all aimed at one demanding target: the receptor-binding domain of the SARS-CoV-2 spike protein, the piece of the virus that latches onto human cells. In the first task, participants were handed early screening data for a "parent" antibody and asked to design new versions that bound the target more tightly while remaining manufacturable. The second asked them to rank a set of existing candidates from strongest to weakest binder. The third asked them to invent entirely new binding loops the screening data had never contained.
The results split cleanly. On the design task, several groups produced strong, developable antibodies, and the best of them beat the answer a wet lab had reached through months of physical experiments. On the other two tasks, the models largely failed. As the authors report, apart from a single model, ranking candidate antibodies "was worse than random clone picking," and inventing new loops from scratch was "highly variable," with many entries failing to beat a standard laboratory screen.
The clearest win came from Aureka Biotechnologies, whose model (submitted as AuraIDE and cited in the paper as AuraBind) topped the design task. According to the company, its winning antibody bound the target at 94.7 picomolar. That is roughly a 2,000-fold tighter grip than the parent antibody, and tighter than the 113-picomolar best the organisers had reached through a further round of laboratory screening. A sequence the model proposed in under a week edged out the product of months of bench work.
What an antibody is, and why binding strength is the whole game#
An antibody is a Y-shaped protein the immune system uses to recognise threats. The tips of the Y carry small, highly variable loops called complementarity-determining regions, or CDRs. These are the parts that actually grip a target. One of them, the heavy-chain third loop (HCDR3), does most of the recognising and varies enormously between antibodies. Redesigning these loops is how you change what an antibody sticks to, and how firmly.
That firmness is measured as affinity, usually written as a dissociation constant, or KD. The counter-intuitive part is that lower numbers mean tighter binding, so a 94.7-picomolar antibody clings far more strongly than a nanomolar one, a thousand-fold difference. For a drug, tighter binding often means a lower dose and fewer side effects. The process of nudging a decent antibody towards this kind of grip is called affinity maturation, and in the lab it is slow, iterative and expensive: build a library of variants, screen it, pick the winners, repeat.
There is a second requirement that rarely makes headlines but sinks many promising molecules: developability. An antibody can bind beautifully and still be useless if it clumps together, sticks to the wrong things or falls apart when warmed. AIntibody scored every design on five such properties (hydrophobicity, polyreactivity, self-interaction, thermal stability and aggregation), and only molecules that cleared all of them counted. That detail matters, because a model that chases affinity while ignoring manufacturability is optimising for the wrong prize.
Finally, the word "blinded" is doing real work here. The design is borrowed from CASP, the structure-prediction contest that ran for years before AlphaFold2 stunned the field in 2020 by predicting protein shapes almost as well as experiments could measure them. CASP's power was that entrants were judged on brand-new structures nobody had seen, so a model could not simply memorise the answer. AIntibody applies the same discipline to antibodies: prospective, independent and scored at the bench.
Why this is important#
Antibodies are not a niche. They are among the best-selling and fastest-growing classes of medicine, with roughly 140 antibody-based drugs approved and annual sales in the region of several hundred billion dollars, spanning cancer, autoimmune disease and infection. Anything that shortens the years and the tens of millions of dollars it takes to mature one candidate has an outsized effect on what gets made, and how quickly patients see it.
The deeper significance is not the winning number but the existence of an honest scoreboard. Generative AI is already woven through antibody discovery, yet, as the authors note, the field has had no agreed way to tell a genuinely capable model from a well-marketed one. Companies tend to grade their own homework, reporting retrospective scores on their own data. A blinded, wet-lab-anchored benchmark strips out that self-flattery. It is the difference between a company claiming its model works and a referee confirming it under controlled conditions.
For medicine and pandemic preparedness, the design result is a genuine, if narrow, encouragement. The task the models won, improving an existing hit, is exactly the slow step in turning a promising antibody into a therapy. If AI can compress rounds of maturation into a week of computation, the response to the next emerging pathogen could look very different.
Critical analysis#
The strengths of this study are structural, not rhetorical. Blinding, independent expression of every design, uniform measurement and hard developability filters make the affinity-maturation result difficult to wave away. When a model designs an antibody that beats months of bench work and survives five manufacturability tests, something real has happened.
The limitations are just as clear, and the paper is refreshingly candid about them. First, the win was confined to one of three tasks. The same models that excelled at maturation were, with one exception, worse than random at ranking candidates and erratic at designing new loops. Success in one regime did not carry over to the others, which is precisely what you would not want if you were hoping AI had "solved" antibody design.
Second, the task the models won was the one richest in relevant data. Challenge 1 handed participants real screening output from the parent antibody's first maturation stage, so the models were extrapolating one careful step beyond high-quality experimental data, not conjuring binders from nothing. That is useful, but it is a long way from the fully computational, target-to-drug pipeline sometimes implied by press releases.
Third, everything here rests on a single antigen. The spike RBD is among the most heavily studied proteins on Earth, with abundant public data. That makes it a fair test in some ways and a soft one in others. Whether these results hold for a novel target with little prior information is the open question, and the one that matters most for real drug programmes.
A note of caution on the headline figure: the 94.7-picomolar result and the "2,000-fold" framing come from the winning company's own announcement of a peer-reviewed result. The independent, blinded testing lends it weight the usual corporate claim lacks, but readers should keep the source in view. As for timelines, this is a benchmark, not a medicine. None of these antibodies is a drug, and the path from a tight-binding, developable sequence to an approved therapy still runs through years of preclinical and clinical work.
An early CASP moment, not an AlphaFold one#
The natural comparison is AlphaFold, and it is instructive mainly for how it differs. AlphaFold predicts a static structure: where the atoms sit. Designing an antibody is a harder, messier problem because you must generate a new sequence that folds, binds a specific target, and is sufficiently stable to manufacture. It is closer to a design competition than a prediction one, which is why a blinded, wet-lab benchmark matters so much more than another leaderboard scored in silico.
AIntibody also differs from the de novo binder work that has drawn attention from groups such as the Baker lab at the University of Washington and Google DeepMind's AlphaProteo. Those methods design proteins to grip a target from scratch and are validated in-house by their creators. AIntibody's contribution is not a new model but a shared, independent yardstick: the connective tissue a field needs to turn scattered claims into cumulative progress. If it becomes for antibodies what CASP became for structure prediction, its most important output will not be this year's winner but the honest measurement of where the models stand each year.
Seen that way, the fairest summary is neither "AI beats the lab" nor "AI flunks the test." It is that AI has crossed a real threshold in one well-defined, data-rich task, while the broader problem of designing antibodies to order remains unsolved. For the first time, though, it is being measured in a way everyone can trust.
Key takeaways#
- A blinded benchmark has arrived. AIntibody is the first international contest to prospectively test AI-designed antibodies, with every entry built and measured in a single laboratory under identical conditions.
- AI won the maturation task. On improving an existing antibody, the best model produced a developable design binding at 94.7 picomolar, beating the best result from months of bench screening.
- It lost the other two. Ranking candidates and inventing new binding loops mostly did no better than chance or a standard lab screen. Skill did not transfer across tasks.
- Data richness was decisive. The models excelled where they had strong experimental data to extrapolate from, on one of the most-studied proteins in biology.
- This is a yardstick, not a cure. The value lies in trustworthy, independent measurement of what the models can really do, not in any single antibody, none of which is yet a drug.
Frequently asked questions#
What is AIntibody? A blinded, international benchmark, published in Nature Biotechnology on 19 August 2026, that tested 511 AI-designed or AI-ranked antibodies from 29 organisations. All were validated in the same laboratory, with results hidden from participants until the contest closed.
Did AI really beat a laboratory? On one task, yes. The winning model designed an antibody that bound more tightly (94.7 picomolar) than the best candidate the organisers found through an extra round of physical screening (113 picomolar). This applied only to the affinity-maturation task, not the other two.
Does this mean AI can now design any antibody? No. The models succeeded at improving an existing antibody with rich supporting data, but were mostly no better than random at ranking candidates or designing new binding loops. General-purpose antibody design remains unsolved.
Is this the same as AlphaFold? No. AlphaFold predicts protein structure. AIntibody tests antibody design, which means generating new, functional, manufacturable sequences. That is a harder problem and, unlike a structure prediction, must be checked experimentally.
Are any of these antibodies going to be medicines? Not yet. The benchmark measured design capability, not therapeutic readiness. Any candidate would still face years of preclinical and clinical testing before it could become a drug.
Why should I trust the winning numbers? Because they came from blinded, independent wet-lab testing rather than a model grading its own work. The headline affinity figure is from the winner's own announcement, but it describes a peer-reviewed, independently validated result.
Glossary#
Antibody: A Y-shaped immune protein that recognises and binds a specific target, such as part of a virus.
CDR (complementarity-determining region): The variable loops at the tips of an antibody that grip its target. HCDR3 is the most variable and important of them.
Affinity (KD): A measure of binding strength. Lower values mean tighter binding; picomolar is very tight, and nanomolar is a thousand-fold weaker.
Affinity maturation: The process of improving an antibody's binding strength, traditionally through repeated rounds of laboratory screening.
Developability: Whether an antibody can actually be manufactured and used, covering stability, solubility and a tendency not to clump or stick to the wrong things.
Blinded benchmark: A test in which entrants are scored on unseen problems, so they cannot memorise or reverse-engineer the answers.
Receptor-binding domain (RBD): The part of the SARS-CoV-2 spike protein that attaches to human cells. It was the target used throughout this contest.
References#
- Erasmus, M. F. et al. "A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability." Nature Biotechnology, 19 August 2026. Peer-reviewed. https://www.nature.com/articles/s41587-026-03238-6
- Erasmus, M. F. et al. "AIntibody: an experimentally validated in silico antibody discovery design challenge." Nature Biotechnology, 2024. Peer-reviewed. https://www.nature.com/articles/s41587-024-02469-9
- Aureka Biotechnologies. "Nature Biotechnology | Aureka Wins the Global Blinded AI Antibody Benchmark." PR Newswire, 26 August 2026. Company announcement. https://www.prnewswire.com/news-releases/nature-biotechnology--aureka-wins-the-global-blinded-ai-antibody-benchmark-ai-design-surpasses-the-best-experimental-result-302860936.html
- "AI designs new antibodies that pass blinded laboratory tests." Phys.org, 27 August 2026. Science journalism. https://phys.org/news/2026-08-ai-antibodies-laboratory.html
- Jumper, J. et al. "Highly accurate protein structure prediction with AlphaFold." Nature, 2021. Peer-reviewed. https://www.nature.com/articles/s41586-021-03819-2
- "The Therapeutic Monoclonal Antibody Product Market." BioProcess International. Industry analysis. https://www.bioprocessintl.com/economics/the-therapeutic-monoclonal-antibody-product-market
- Google DeepMind. "AlphaProteo generates novel proteins for biology and medicine." Research lab announcement. https://deepmind.google/blog/alphaproteo-generates-novel-proteins-for-biology-and-health-research/