The night shift nobody sees#

Picture a laboratory at 3am. The benches are empty, the lights are low, and somewhere in a data centre an AI agent is still reading papers, proposing an experiment, and queuing up the analysis it wants to run when the humans come back. That scene stopped being science fiction some time ago. What changed this week is the plumbing underneath it.

On 29 September 2026, OpenAI used its annual DevDay event to push autonomous agents from a demo into standard infrastructure. The company introduced "always-on" agents it calls dots, upgraded its Agents API so agents can operate software and run sub-agents in parallel, and worked with Amazon on Bedrock Managed Agents so those agents can run inside a company's own cloud. None of it is aimed specifically at science. That is exactly why it matters. The bottleneck for automating research has rarely been a clever idea. It has been the boring machinery that keeps an agent running for hours, remembers what it did, uses real tools, and does not fall over halfway through.

This piece is about that shift and what it does, and does not, mean for the people who actually do science.

Background: what "automating research" really means#

When people say AI is automating research, they usually picture a chatbot writing a literature review. The more interesting work is broader. Science runs on a loop: read what is known, form a hypothesis, design an experiment, run it, analyse the results, then update the hypothesis and go again. For decades, AI could help with one slice of that loop at a time. Automating the whole cycle, end to end, stayed out of reach.

A few terms are worth pinning down before we go further.

An AI agent is a program built on a large language model that can plan a task, call tools such as search engines or code interpreters, act on the results, and keep going without a human prompting each step. A large language model (LLM) is the underlying system trained on huge amounts of text to predict and generate language, which is what gives an agent its reasoning and writing ability. A multi-agent system is several of these agents working together, often with different jobs, so one might generate hypotheses while another critiques them and a third runs the numbers. A self-driving lab, sometimes called an autonomous laboratory, pairs that software with real robots, so the agent can physically run experiments and read the results back into its next decision. And agentic AI is the umbrella term for all of this: AI that does things rather than just answering questions.

Keep the research loop in mind, because the story of 2026 is really the story of AI creeping around more and more of it.

Why DevDay 2026 is an infrastructure moment, not a science one#

The headline features from DevDay sound like productivity tools, and for most users they are. Dig into the developer announcements, though, and you find the parts that matter for research automation.

The updated Agents API now supports "computer use," multi-agent workflows, tool search, and context compaction, with OpenAI running the underlying infrastructure. In plain terms: an agent can now click around real software, spin up helper agents, search through a large set of tools, and keep track of a long task without running out of memory, while someone else keeps the servers alive. Those are precisely the capabilities a research agent needs to operate a lab information system, a data pipeline, or an instrument dashboard for hours at a time.

The dots agents are designed to run continuously on your behalf rather than in a single chat session, and Bedrock Managed Agents let organisations run these agents inside their own AWS environment. For a pharmaceutical company or a university lab, that last point is not a detail. Data governance and keeping sensitive results in-house are often the difference between a tool you can use and one you cannot.

So the significance is indirect but real. DevDay did not announce a discovery. It lowered the cost and effort of building the kind of persistent, tool-using agent that science automation has been waiting for. When the scaffolding gets cheaper, more people build on it.

The proof it works: AI scientists are already publishing#

Here is the part that surprises people. While the infrastructure was maturing, the science-specific systems quietly cleared some genuine bars.

Earlier in 2026, Nature published "Towards end-to-end automation of AI research", describing a system that generates research ideas, writes the code, runs the experiments, analyses the data, and drafts the full manuscript, then reviews its own work. One paper it produced passed the first round of peer review at a workshop of a major machine-learning conference. That is a narrow result, a workshop rather than a flagship journal, and the authors are careful about it. It is still the first time an automated pipeline navigated the whole loop and had the output judged, blind, by human reviewers.

Then there is Robin. A separate Nature paper described a multi-agent system that automates both hypothesis generation and data analysis for experimental biology. Working on dry age-related macular degeneration, an eye disease with limited treatments, the system proposed boosting a clean-up process in retinal cells and flagged two existing compounds, ripasudil and KL001, which were then confirmed to work in cell experiments. The humans still ran the wet-lab validation. The hypotheses that pointed them there came from the machine.

Commercial systems are pushing the same direction. Edison Scientific's Kosmos is a hosted "AI scientist" that runs self-directed research campaigns and produces a traceable report. Its technical report, a preprint that has not been peer reviewed, describes seven findings, three that reproduced known results and four described as new contributions to the literature, including statistical evidence linking a protein called SOD2 to reduced heart-tissue scarring. Independent scientists quoted by GEN estimated that one long Kosmos run can compress roughly six months of PhD-level work into a single day. Treat that number as a vendor-adjacent claim, not a settled fact, but the direction is clear.

The data problem, and a deal that speaks volumes#

An AI scientist is only as good as what it reads, and what it reads is skewed. Marinka Zitnik of Harvard Medical School put the problem bluntly: about 95% of life-sciences publications focus on only around 5,000 of the most-studied human genes. An agent that only reads the literature inherits that bias and keeps proposing hypotheses about the genes everyone already studies.

That is why a quiet piece of business news is more important than it looks. In September 2026, Edison Scientific signed an agreement with Springer Nature to bring licensed content from 60 specialist journals across the Nature Portfolio into Kosmos. The race is shifting from who has the cleverest agent to who can legally feed it the best, broadest, most trustworthy data. Licensing deals, not model sizes, may end up shaping which systems generate reliable hypotheses.

When the agent touches the real world#

Software agents can reason about experiments. Self-driving labs actually run them. This is where automation stops being about text and starts changing the bench.

At the US Department of Energy's Oak Ridge National Laboratory, a team used AI to guide data collection in real time while mapping residual strain inside a 3D-printed metal part used in turbines and aerospace, working at Cornell's high-energy synchrotron. Instead of collecting data on a fixed schedule and analysing it later, the system decided what to measure next as the experiment unfolded. In chemistry and materials, self-driving labs that switch from batch experiments to continuous, AI-directed ones have been reported to collect around ten times more data, which speeds up the search for new materials.

Close that loop and the picture from the top of this article stops being a metaphor. The reading, the hypothesis, the experiment, and the analysis all start to happen without a person driving each step. The community is organising around this too. A dedicated NeurIPS 2026 workshop on end-to-end autonomous research with robot and AI scientists is scheduled for December, covering everything from literature synthesis to physical experimentation and peer review.

The honest limits#

It would be easy to end on a triumphant note. The reality is messier, and the researchers building these systems are often the most cautious voices in the room.

Start with reproducibility, which was a problem long before AI arrived. A widely cited Nature survey found that more than 70% of researchers had tried and failed to reproduce another scientist's experiments, and more than half had failed to reproduce their own. Automation can help by logging every step an agent takes, but it can also flood the literature with plausible-looking results that nobody has checked. Speed cuts both ways.

Then there is the gap between a hypothesis and the truth. As one biologist put it, AI can search enormous datasets and generate ideas, but a scientist still has to decide whether an idea is biologically meaningful and whether it holds up in a real experiment. In the Robin work, the machine proposed and the humans confirmed. That division of labour is the current state of the art, not a temporary inconvenience.

And the flashy numbers deserve scrutiny. "Six months of work in a day" and "ten times more data" are real signals, but they come from specific setups, often reported by the teams with something to gain. The manuscript that passed peer review passed at a workshop, not at Nature itself. None of this is a reason to dismiss the progress. It is a reason to read the fine print, which is the whole point of good science anyway.

Key takeaways#

The most useful way to hold all of this: infrastructure and capability advanced on separate tracks in 2026, and they are now converging.

First, DevDay 2026 matters because it made long-running, tool-using agents into ordinary infrastructure, which lowers the barrier for anyone building research automation on top.

Second, AI scientists have crossed real thresholds, from a manuscript passing workshop peer review to a multi-agent system surfacing drug candidates later confirmed in the lab.

Third, data access is becoming the real battleground, and licensing deals like Edison Scientific and Springer Nature signal where the advantage will sit.

Fourth, self-driving labs are closing the loop between reasoning and physical experiments, especially in chemistry, materials, and parts of biology.

Fifth, humans remain in the loop for judgement and validation, and the reproducibility and hype problems are real enough that scepticism is a feature, not a flaw.

Frequently asked questions#

Is AI going to replace scientists? Not on current evidence. The systems that work best keep humans in charge of judgement and physical validation. A fairer summary is that scientists who use these tools may out-compete those who do not, rather than being replaced by the tools themselves.

What did OpenAI actually announce for research at DevDay 2026? Nothing science-specific. The relevant changes were general: always-on dots agents, an Agents API that can operate software and run parallel sub-agents, and managed agents that run inside a company's own cloud. Those capabilities make research agents easier to build.

Has an AI really made a scientific discovery? It depends on your bar. A multi-agent system proposed drug candidates for an eye disease that were then confirmed in cell experiments, with humans doing the validation. Fully independent, unsupervised discovery has not happened.

Are these AI-written papers peer reviewed? Some are, some are not. The end-to-end system's paper passed a workshop's first review round. By contrast, the Kosmos technical report is a preprint and has not been peer reviewed. Always check which is which.

What is a self-driving lab? A laboratory where AI plans experiments and robots carry them out, feeding results back so the AI decides what to do next. Oak Ridge National Laboratory has used AI to steer data collection in real time.

What is the biggest risk? Volume without verification. If automation floods journals with unchecked results, it could worsen the reproducibility problems the field already has.

How can a researcher start using this now? Begin with narrow, checkable tasks such as literature synthesis or data analysis, keep a human reviewing every output, and favour tools that log their steps so you can audit what the agent actually did.

Glossary#

AI agent: A program built on a language model that can plan a task, use tools such as search or code, act on the results, and continue without a human prompting every step.

Large language model (LLM): An AI system trained on large amounts of text to generate and reason with language. It is the engine underneath most agents.

Multi-agent system: Several AI agents working together on different parts of a task, for example one generating hypotheses and another checking them.

Agentic AI: A general term for AI that takes actions and pursues goals over time, rather than only answering single questions.

Self-driving lab: An autonomous laboratory where AI designs experiments and robots run them, with results feeding back into the AI's next decision.

Hypothesis generation: The step where a system proposes a testable explanation or prediction, which then has to be checked by experiment.

Reproducibility: Whether an experiment gives the same result when repeated by others. A long-standing weak point in science that automation could help or harm.

Peer review: The process where independent experts evaluate a study before publication. Passing it at a workshop is a lower bar than passing it at a flagship journal.

References#