AI
AI Co-Scientists Are Here, and Nature Just Put Them to the Test
A Nature feature published this week tracks how AI co-scientist platforms from Google DeepMind, FutureHouse and others are already generating validated hypotheses in real labs, and why human judgement is becoming the scarcest resource in science.
A biochemist at the Whitehead Institute typed a question into an AI system, hit enter, and went off to do a long-distance triathlon. By the time she was back, the system had read more than 700 papers and handed her a drug strategy nobody on her team had considered. That is not a promotional anecdote from a startup's blog. It is the opening scene of a Nature Technology Feature published on 21 September 2026, and it captures, better than any press release could, where AI-assisted science actually stands right now: genuinely useful, occasionally strange, and still entirely dependent on a human deciding what counts as a good answer.
Science journalist Elie Dolgin's feature in Nature pulls together the clearest picture yet of how "AI co-scientist" platforms are being used inside working labs, rather than in benchmark tests. The centrepiece is Google DeepMind's Co-Scientist, a multi-agent system built on the Gemini model family that Whitehead Institute researchers Anna Pertl and Kalon Overholt used to attack one of cancer biology's toughest targets: the protein MYC, which drives runaway growth in most cancers and has resisted direct drug targeting for decades.
After the researchers spent close to an hour correcting the system's misunderstanding of their question, Co-Scientist combed through more than 700 papers, generated 108 candidate strategies, and converged on one: rather than dissolving the molecular condensates that cluster around the MYC gene, glue them together using click chemistry, jamming the gene-reading machinery entirely. Overholt called the idea "extremely conceptually compelling," adding that the team had "never thought about anything like this."
Co-Scientist's track record extends beyond that one case. In a study published in Nature in May 2026, a Google-led team led by Juraj Gottweis used the system to propose drug-repurposing candidates for acute myeloid leukaemia (AML). Independent laboratory testing found that one candidate, binimetinib, a drug already approved for melanoma, killed AML cells at nanomolar concentrations. A rival platform, Robin, built by the non-profit lab FutureHouse, produced a similarly validated result for an eye disease published the same month.
Perhaps the most striking case in Dolgin's feature involves Imperial College London microbiologists José Penadés and Tiago Costa, who study how bacteria swap DNA across species. The pair had already spent years collecting experimental data and forming a theory, but for a controlled test, they withheld their findings and asked Co-Scientist to work from the public literature alone. Working only from published papers, the system reached essentially the same conclusion the researchers had, in about two days, a result published in Cell in 2025. Costa's summary: "If you think about scientific discovery as a 100-metre race, with an AI assistant you start this race at, say, metre 30 instead of metre 0."
How an AI Co-Scientist Actually Works#
Unlike a standard chatbot, which answers a prompt in a single pass, systems such as Co-Scientist and Robin use what is called a multi-agent architecture. The AI dispatches several separate reasoning processes, or agents, each pursuing a distinct line of enquiry: one might search the literature, another might generate competing hypotheses, another might critique and refine them. These agents run in parallel, sometimes for many hours, consuming substantial computing power before converging on a small set of workable ideas.
This differs from earlier "AI for science" tools in an important way. Systems like AlphaFold, the protein-structure predictor from Google DeepMind, are narrow specialists: trained to solve one well-defined problem extremely well. A co-scientist platform is closer to a research collaborator that can take an open-ended question, break it into sub-problems, search and reason across disconnected fields, and propose a testable next step. FutureHouse's Kosmos platform, mentioned in the same feature, was used by Stanford and FutureHouse researcher Philine Guckelberger to sift a 30-million-row genomic dataset and surface a cancer-cell pattern she later confirmed independently, work she says "accelerated this like crazy."
Other groups are pursuing the same goal with different designs. Sakana AI's "AI Scientist," described in Nature in April 2026, goes a step further by attempting the entire research cycle: generating an idea, running the experiment through automated code, and drafting the paper. Huawei and academic groups including Stanford's Virtual Lab project, published in Nature in 2025, have built comparable multi-agent systems aimed at giving researchers "instant access to interdisciplinary experts," in the words of Virtual Lab co-developer Kyle Swanson.
The design choice that separates these platforms from a search engine or a chatbot is what happens between the question and the answer. Instead of retrieving the closest match to a prompt, each agent builds and tests a small model of the problem, checks it against the literature, and passes its conclusions to the other agents for critique. That back-and-forth is slow and expensive. It's also, according to the researchers quoted in the Nature feature, the reason the systems occasionally land on ideas nobody had written down before.
Why This Matters#
For biotechnology and drug discovery, the case now rests on wet-lab validation, not simulation. The AML and eye-disease results are the clearest evidence yet that these tools can shorten the distance between a computational hypothesis and a compound that actually works in cells. For public health and pandemic preparedness, the same architecture, agents that can rapidly synthesise scattered literature and propose testable mechanisms, could plausibly compress the early, time-critical phase of investigating an emerging pathogen: working out a resistance mechanism, say, or flagging drugs worth repurposing.
For the scientific workforce, the implications sit less easily. Economist Ajay Agrawal of the University of Toronto, quoted in the Nature piece, argues that as AI absorbs more of the mechanical work of research, "the most valuable part now is actually asking the question." That changes what training a scientist even means. If AI agents can generate and test hypotheses faster than any single lab, the scarce skill shifts from execution to judgement: knowing which questions are worth asking, and which of the machine's 108 suggestions is the one worth funding.
Critical Analysis#
The feature is careful not to overclaim. Every validated success it cites involved a domain expert framing the question precisely, correcting the AI's early misunderstandings, and independently verifying the output before trusting it. Fyodor Urnov of the University of California, Berkeley, an early user of FutureHouse's Kosmos, sums up his working rule with a Cold War arms-control phrase: "trust, but verify." Guckelberger's experience with Kosmos misreading a data column, and the Whitehead team's hour-long back-and-forth correcting Co-Scientist's basic assumptions about condensates, both show these systems still need careful supervision rather than blind faith.
There is also a broader, quantitative reality check. A state-of-the-industry report covered by Nature in April 2026, based on Stanford's AI Index, found that on complex, open-ended tasks, human scientists still substantially outperform the best AI agents, even as the same report noted a nearly 30-fold rise in AI-mentioning science publications between 2010 and 2025. The gap between "wins in a curated demonstration" and "reliably outperforms a postdoc on a genuinely novel problem" remains real.
Timeline-wise, this is incremental rather than revolutionary. Co-Scientist and Robin have each produced a small number of validated results, not a systematic replacement for hypothesis-driven research. Running dozens of agents through hundreds of papers for a single question is computationally expensive, and nobody quoted in the feature offers a figure for what that costs per validated idea. Reproducibility across labs is also untested at scale: a result that emerges from one lab's specific framing of a question may not reappear if a different team asks it slightly differently. And the risk that matters most for anyone outside the small group of experts currently using these tools is quieter than a wrong answer: a confidently stated, plausible-sounding error that nobody catches because the reviewer trusted the system a little too much.
Expert Perspective#
What distinguishes this moment from earlier "AI for science" milestones like AlphaFold is scope. AlphaFold solved a single, extremely well-specified problem, protein structure prediction, to a degree that reshaped structural biology. Co-Scientist and its rivals aim instead at the open-ended, messy front end of research: deciding what to investigate and how. That is a fundamentally harder target, and the Nature feature's own reporting shows it is only partially met. Le Cong of Stanford, a co-founder of the rival platform Phylo, frames the ambition modestly: "We're trying to move humans up the value chain," not replace them. Hector Zenil of King's College London goes further, arguing that researchers with strong human judgement "are going to be more valuable than ever," precisely because AI is automating the parts of science that used to train that judgement in junior researchers.
Key Takeaways#
Validated wet-lab results, not benchmark scores, are what separate this generation of AI co-scientist tools from earlier hype. Multi-agent architectures let these systems explore many hypotheses in parallel rather than reasoning down a single chain. Human framing of the question and independent verification of the output remain essential at every step so far demonstrated. A parallel report shows human scientists still outperform AI agents on complex, open-ended tasks, tempering the more dramatic claims. The bottleneck in science may be shifting from execution capacity to the human skill of asking sharp, well-scoped questions.
Frequently Asked Questions#
What is an AI co-scientist? It is a multi-agent AI system, such as Google DeepMind's Co-Scientist or FutureHouse's Robin, designed to help generate hypotheses, search literature, design experiments and analyse data, working alongside a human researcher rather than autonomously.
Has an AI co-scientist actually discovered a new drug? Not independently. Co-Scientist proposed drug-repurposing candidates for acute myeloid leukaemia that researchers then tested in the lab, where one, binimetinib, showed strong activity against leukaemia cells at nanomolar concentrations, published in Nature in May 2026.
Is this the same technology as AlphaFold? No. AlphaFold predicts protein structure, a single well-defined task. Co-Scientist-style systems tackle open-ended reasoning across an entire research question, a broader and less mature capability.
Do these tools replace human scientists? According to every researcher quoted in the Nature feature, no. Each success required a scientist to frame the question, correct early misunderstandings, and independently verify the results.
Which organisations are building these systems? Google DeepMind (Co-Scientist), FutureHouse (Robin and Kosmos), Sakana AI (AI Scientist), Huawei, and academic groups including Stanford's Virtual Lab and the start-up Phylo.
How reliable are AI co-scientist outputs? Reliability varies by task and requires expert verification. A parallel Stanford AI Index report found human scientists still outperform the best AI agents on complex tasks, even as adoption of AI in science has grown sharply.
Are these systems available to any researcher? Some, including FutureHouse's Kosmos, are accessible to academic users; access and cost vary by platform and are evolving quickly as of September 2026.
What is the biggest current limitation? The need for precise question-framing and independent verification. These systems can misinterpret data or assumptions, as seen when Co-Scientist initially misunderstood the Whitehead team's condensate question and when Kosmos misread a genomic data column.
Glossary#
AI co-scientist: A multi-agent AI system designed to assist scientific research by generating hypotheses, searching literature and helping design experiments.
Multi-agent system: An AI architecture in which several autonomous reasoning processes ("agents") work on different aspects of a problem in parallel, then converge on a shared output.
Condensate: A liquid-like cluster of proteins and DNA that forms inside cells and can concentrate the molecular machinery needed to switch genes on or off.
Super-enhancer: A cluster of DNA control regions that strongly boosts the activity of nearby genes, including cancer-driving genes such as MYC.
Click chemistry: A class of chemical reactions designed to join molecules together quickly, selectively and reliably, widely used in drug and materials design.
Drug repurposing: Identifying new therapeutic uses for medicines that are already approved for a different condition.
IC50 (half-maximal inhibitory concentration): A standard laboratory measure of how much of a compound is needed to reduce a biological process, such as cell survival, by half; lower values indicate greater potency.
AI Index: An annual state-of-the-field report from Stanford's Institute for Human-Centered AI tracking trends in AI research, investment and capability.
References#
- Dolgin, E. "AI co-scientists are revolutionizing how research is done." Nature 657, 1108–1110 (2026). https://www.nature.com/articles/d41586-026-02931-5
- Gottweis, J. et al. "Accelerating scientific breakthroughs with an AI co-scientist." Nature 655, 487–496 (2026). https://doi.org/10.1038/s41586-026-10644-y
- Ghareeb, A. E. et al. Nature 655, 497–505 (2026). https://doi.org/10.1038/s41586-026-10652-y
- Penadés, J. R. et al. Cell 188, 6654–6665 (2025). https://doi.org/10.1016/j.cell.2025.08.018
- Lu, C. et al. "The AI Scientist." Nature 651, 914–919 (2026). https://doi.org/10.1038/s41586-026-10265-5
- Swanson, K., Wu, W., Bulaong, N. L., Pak, J. E. & Zou, J. "Virtual Lab." Nature 646, 716–723 (2025). https://doi.org/10.1038/s41586-025-09442-9
- Jones, N. "Human scientists trounce the best AI agents on complex tasks." Nature 652, 841–842 (2026). https://www.nature.com/articles/d41586-026-01199-z
- Stanford Institute for Human-Centered AI. Artificial Intelligence Index Report 2026. https://hai.stanford.edu/ai-index/2026-ai-index-report
- "Why AI cannot do good science without humans." Nature. https://www.nature.com/articles/d41586-026-01551-3