A lab meeting with almost no humans in the room#
Picture a lab meeting. A principal investigator opens with a question about a fast-mutating virus. A biologist, a computational scientist and a sceptical critic each take a turn. They argue, agree a plan, and write the code to carry it out. The whole discussion takes an hour or two. Every voice in the room is a language model.
That is not a thought experiment. In a study published in Nature, a "Virtual Lab" of AI agents designed 92 new nanobodies (small antibody-like proteins) against variants of the virus that causes COVID-19, and two of them showed improved binding to the recent JN.1 or KP.3 variants while still binding the ancestral virus strongly (Swanson et al., Nature). A human researcher gave high-level feedback, and the results were checked in a real wet lab.
This is the idea behind multi-agent AI for science: instead of one chatbot answering questions, you build a team of specialised AI roles that propose, criticise, rank and test ideas. The approach is generating real results and real hype. This post separates the two. We will look at how these systems are built, what they have actually delivered at the bench, and the ways they fail that their own creators now document.
Background: what a multi-agent AI lab actually is#
A large language model (LLM) is a program trained on enormous amounts of text to predict what comes next. Used alone, it answers a question and stops. An agent is an LLM given extra abilities: it can search the literature, run code, call other software, or operate lab equipment, and it can decide its own next step. A multi agent system puts several of these agents in one workflow, each with a defined job, much like a research group where one person reads papers, another runs experiments and a third plays devil's advocate.
Why split the work at all? One reason is focus. A model told to act only as a critic behaves differently from the same model told to be creative. Another is checking. When one agent's output becomes another agent's input, mistakes have a chance of being caught. A review of LLM-based hypothesis generation groups the field into four approaches (direct prompting, retrieval-enhanced methods, multi-agent teams that simulate research groups, and reasoning-focused systems) and notes that evaluating them remains hard, with hallucination and the balance between novelty and feasibility as persistent problems (Herron et al., ACM Computing Surveys).
Two further terms will recur. Test-time compute means letting a model think for longer, or run many attempts, when answering. A self-driving laboratory is a lab where robots carry out experiments and software chooses what to try next. Combine the two ideas and you get what some call an autonomous lab: agents that plan, robots that execute, and results that feed back into the next plan. How close anyone has come to that, and how much of it is demonstration rather than routine practice, is what the rest of this post examines.
Co-Scientist: a tournament for ideas#
Google's Co-Scientist is the best-documented example of multi-agent AI applied to hypothesis generation. It is built on the Gemini model family and is designed so that agents continuously generate, critique and refine hypotheses, with a "tournament evolution" process that makes the hypotheses improve over time (Gottweis et al., Nature).
The cast, as Google describes it, is easy to follow. A Generation agent proposes hypotheses from the literature and data, and a Proximity agent clusters them so the system does not keep exploring the same idea. A Reflection agent acts as a virtual peer reviewer. A Ranking agent runs a pairwise "idea tournament" in which hypotheses debate each other. An Evolution agent refines and combines the winners, a Meta-review agent writes up the result, and a supervisor agent schedules all of it in parallel (Google DeepMind).
The system's own automated evaluations show that hypothesis quality keeps improving as more test-time compute is spent (Gottweis et al., Nature). That is a useful property but also a caution, because the judge in these evaluations is itself a model.
The stronger evidence comes from the bench. The authors report that Co-Scientist helped identify drug-repurposing candidates and synergistic combination therapies for acute myeloid leukaemia that were validated in laboratory (in vitro) experiments (Gottweis et al., Nature). Google's announcement adds collaborator case studies, including a liver fibrosis candidate that blocked 91% of a scarring-linked response in lab tests (Google DeepMind). Treat that second set with care: those figures are company-reported summaries, and the page itself gives few methodological details.
Smaller academic teams have reported the same pattern. A framework called Coated-LLM, with Researcher, Reviewer and Moderator agents, predicted effective drug combinations for Alzheimer's disease with an accuracy of 0.74 against 0.52 for a traditional knowledge-based method, and one predicted combination reduced amyloid aggregation in vitro (Xu et al., iScience).
The Virtual Lab: a principal investigator made of text#
If Co-Scientist is a tournament, the Virtual Lab is a meeting. An LLM principal investigator agent guides a team of scientist agents through a series of research meetings, with a human providing feedback. In the nanobody project, the agents built a design pipeline that combined the protein language model ESM, the structure predictor AlphaFold-Multimer and the modelling software Rosetta (Swanson et al., Nature).
A protein language model, for readers new to the term, reads amino-acid sequences the way a chatbot reads sentences and learns which patterns are plausible. AlphaFold-Multimer predicts how protein chains fit together. Rosetta scores the physical fit. The agents did not invent these tools. They chose them, wired them together and wrote the code, which is the part a human postdoc would normally spend weeks on.
The Stanford team behind the work told the student newspaper that the core planning took one to two hours of agent discussion, and that design took days rather than the weeks or months a human researcher might need (Stanford Daily). The same report is candid about the limits. The agents do not know a lab's real equipment or interests, they can suggest impractical experiments, and one of the researchers said they were too agreeable and did not challenge each other enough. They also cannot yet run wet-lab experiments themselves (Stanford Daily).
Notice the arithmetic: 92 designs, two with improved binding to the newer variants. That is a promising hit rate for early-stage design, and it also shows why experimental validation is not optional.
When the agents get hands: chemistry robots and self-driving labs#
Proposing a hypothesis is one thing. Running the experiment is another. Chemistry has been the testing ground for agents with hands, because its reactions are fast, measurable and amenable to automation.
The pioneering example, Coscientist (a different system from Google's Co-Scientist, despite the near-identical name), used GPT-4 to autonomously design, plan and perform experiments by combining internet and documentation search, code execution and experimental automation, and it succeeded in optimising palladium-catalysed cross-coupling reactions (Boiko et al., Nature). Later systems adopted a team structure. LLM-RDF uses six specialised agents, including a Literature Scouter, an Experiment Designer and a Spectrum Analyzer, and guided an alcohol-oxidation reaction from literature search through scale-up and purification (Ruan et al., Nature Communications). ChemAgents puts a Task Manager in charge of four role-specific agents and a robotic lab (Song et al., JACS).
The table below compares the main systems discussed in this post.
| System | Roles (examples) | What it does | Evidence so far |
|---|---|---|---|
| Co-Scientist (Google) | Generation, Reflection, Ranking, Evolution, Meta-review | Generates and ranks hypotheses | Leukaemia drug-repurposing candidates validated in vitro (Nature) |
| Virtual Lab (Stanford) | Principal investigator, specialist scientists, critic | Designs proteins via ESM, AlphaFold-Multimer, Rosetta | 92 nanobodies designed, two with improved variant binding (Nature) |
| Coscientist | Planner plus tool-using modules | Plans and runs chemistry experiments | Cross-coupling optimisation (Nature) |
| ChemAgents | Task Manager, Literature Reader, Experiment Designer, Computation Performer, Robot Operator | Runs multistep robotic chemistry | Six experimental tasks, plus a seventh in a new robotic lab (JACS) |
| The AI Scientist | Ideation, coding, writing, reviewing | Produces a full paper end to end | One manuscript passed workshop review (Nature) |
A 2026 perspective in JACS Au adds a note of discipline: agent-enabled self-driving labs "should not be conflated with autonomous scientific discovery," and the authors list safety in physical execution, hardware interoperability, reproducibility and auditability as central unsolved challenges (Xie et al., JACS Au). In plain terms, a robot that runs your reaction faster has not necessarily had a good idea.
Fully automated papers: the AI Scientist and its critics#
The boldest version of the idea removes the human altogether. The AI Scientist creates research ideas, writes code, runs experiments, analyses data, writes the manuscript and performs its own peer review. In a 2026 Nature paper, its developers reported that one generated manuscript passed the first round of peer review at a workshop of a top-tier machine learning conference, a venue with a 70% acceptance rate. The authors themselves flag risks, including overwhelmed review systems and noise in the literature (Yamada et al., Nature).
Independent testing of an earlier version was less flattering. Beel and colleagues found that five of twelve proposed experiments (42%) failed because of coding errors, that some ideas flagged as novel were well-established concepts, and that manuscripts cited a median of just five papers. They also noted that a full paper cost only $6 to $15 to generate, with about 3.5 hours of human involvement (Beel et al., ACM SIGIR Forum). Cheap and fast, then, but not yet reliable.
A preprint (not peer reviewed) from Luo and colleagues goes further. It identifies four failure modes in open-source AI scientist systems: inappropriate benchmark selection, data leakage, metric misuse and post-hoc selection bias. It also finds that access to trace logs and code makes these failures far easier to detect than reading the final paper alone (Luo et al., preprint, not peer reviewed). The authors recommend that journals require those artefacts alongside AI-generated research. That is a sensible rule for any journal editor to adopt now.

Why teams of agents fail, and how to catch it#
Teams of agents are not automatically smarter than one agent, any more than a committee is automatically wiser than an individual. Several recent studies, mostly preprints, document how multi-agent systems go wrong. Treat these as early signals rather than settled findings.
A taxonomy built from more than 1,600 annotated traces across seven frameworks identified 14 failure modes in three groups: system design problems, misalignment between agents, and weak task verification (Cemri et al., preprint, not peer reviewed). A separate benchmark called HiddenBench tested what happens when each agent holds only part of the evidence. Multi-agent LLMs reached only 30.1% accuracy, against 80.7% for a single agent given all the information, because the agents failed to reason about what others might know and settled too quickly on shared evidence (Li et al., preprint, not peer reviewed). Risk analysts also warn of conformity bias and "monoculture collapse," where agents built on the same model share the same blind spots (Reid et al., preprint, not peer reviewed). That echoes the Stanford team's observation that their agents were too agreeable.
The chart below puts two of the published numbers side by side. They come from different studies measuring different things, so read each panel on its own.
Figure: Left, a multi-agent framework beat a knowledge-based method on a drug-combination task (Xu et al., iScience). Right, a different multi-agent setup did far worse than a single, fully informed agent when evidence was split between agents (Li et al., preprint, not peer reviewed).
There is a safety dimension too. A Nature Communications perspective on the risks of AI scientists argues for prioritising safeguards over autonomy, proposing a framework built on human regulation, agent alignment and understanding of environmental feedback (Tang et al., Nature Communications). Google says it ran safety evaluations of Co-Scientist, including assessments of misuse involving chemical, biological, radiological and nuclear threats, though it gives few details (Google DeepMind).
Where this leaves working scientists#
The honest summary is mixed. Multi-agent AI has produced experimentally tested outputs in more than one field: leukaemia drug candidates, nanobodies, optimised chemical reactions. It has also produced experiments that crashed, ideas that were not new, and agents that nod along with each other.
Three practical lessons follow. First, treat these systems as an AI research assistant that generates leads, not as a source of conclusions. Every hit still needs a bench. Second, keep the logs, because the evidence suggests process records catch errors that finished papers hide (Luo et al., preprint, not peer reviewed). Third, design for disagreement. If your critic agent runs on the same model as your proposer, you may have built an echo chamber with extra steps.
A fair prediction is that the near-term gains will come from speed on narrow, checkable tasks, which is also where reviews of AI agents in biomedicine say humans should stay in the loop as sceptical partners (Gao et al., Cell). The open question is whether open-ended discovery, the kind that surprises its own authors, will follow. Nobody has shown that yet.
Frequently asked questions#
What is a multi-agent AI lab? It is a set of AI agents, each assigned a role such as proposer, critic, coder or robot operator, that work through a research problem together. Systems such as Co-Scientist and the Virtual Lab are built this way (Gottweis et al., Nature).
Are multi-agent systems better than a single AI model? Sometimes. A multi-agent framework beat a knowledge-based method on Alzheimer's drug-combination prediction (Xu et al., iScience), but a benchmark found multi-agent models doing much worse than a single informed agent when information was split between them (Li et al., preprint, not peer reviewed).
Has a multi-agent AI made a real laboratory discovery? There are laboratory-validated outputs, such as leukaemia drug-repurposing candidates from Co-Scientist (Gottweis et al., Nature) and nanobodies from the Virtual Lab (Swanson et al., Nature). Whether any of them becomes a medicine is a separate and much longer question.
Can AI agents run experiments on their own? In chemistry, some can, with robotic hardware. Coscientist planned and performed reactions using tools and experimental automation (Boiko et al., Nature). The Virtual Lab agents, by contrast, could not run wet-lab work themselves (Stanford Daily).
Can an AI write a paper that passes peer review? One AI-generated manuscript passed the first round of review at a machine learning workshop with a 70% acceptance rate (Yamada et al., Nature). Independent testing of an earlier version found frequent experiment failures and thin citations (Beel et al., ACM SIGIR Forum).
What are the main risks? The documented ones include benchmark selection errors, data leakage and metric misuse in AI scientist systems (Luo et al., preprint, not peer reviewed), coordination failures among agents (Cemri et al., preprint, not peer reviewed), and misuse risks that call for safeguards (Tang et al., Nature Communications).
How can researchers try these tools? Google says its Hypothesis Generation tool is experimental and rolling out, with registration available through Google Labs (Google DeepMind). Several academic systems also publish code.
References#
- Gottweis J. et al. Accelerating scientific discovery with Co-Scientist. Nature. Link
- Google DeepMind. Co-Scientist: a multi-agent AI partner to accelerate research (19 May 2026). Link (company announcement; results are company-reported)
- Swanson K. et al. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature (2025). Link
- Stanford Daily. Researchers develop AI scientists for therapeutic discovery (6 May 2026). Link (science journalism)
- Boiko D. A. et al. Autonomous chemical research with large language models. Nature (2023). Link
- Ruan Y. et al. An automatic end-to-end chemical synthesis development platform powered by large language models. Nature Communications (2024). Link
- Song T. et al. A multiagent-driven robotic AI chemist enabling autonomous chemical research on demand. Journal of the American Chemical Society (2025). Link
- Xie Z.-K. et al. Autonomous chemistry and materials innovation driven by scientific agents. JACS Au (2026). Link
- Yamada Y. et al. Towards end-to-end automation of AI research. Nature (2026). Link
- Beel J. et al. Evaluating Sakana's AI Scientist: bold claims, mixed results, and a promising future? ACM SIGIR Forum (2025). Link
- Xu Q. et al. Multi agent large language models for biomedical hypothesis generation in drug combination discovery. iScience (2025). Link
- Herron E. J. et al. From rules to reasoning: a survey of large language model-based approaches to scientific hypothesis and idea generation. ACM Computing Surveys (2026). Link
- Tang X. et al. Risks of AI scientists: prioritizing safeguarding over autonomy. Nature Communications (2025). Link
- Gao S. et al. Empowering biomedical discovery with AI agents. Cell (2024). Link
- Luo Z. et al. The more you automate, the less you see: hidden pitfalls of AI scientist systems (2025). Link (not peer reviewed)
- Cemri M. et al. Why do multi-agent LLM systems fail? (2025). Link (not peer reviewed)
- Li Y. et al. Systematic failures in collective reasoning under distributed information in multi-agent LLMs (2025). Link (not peer reviewed)
- Reid A. et al. Risk analysis techniques for governed LLM-based multi-agent systems (2025). Link (not peer reviewed)