Ageing and Longevity

Tiny AI Models Just Outsmarted the Frontier Giants on Ageing Science

A new Cell study pits 18 leading AI systems against ageing biology data. Small, specialised models beat every frontier system, and the whole toolkit is free to use.

On 17 September 2026, eighteen of the world's most capable AI systems sat the same exam. It covered DNA methylation patterns, blood proteins, and the genetics of animals that age unusually slowly. The systems came from OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI. Not one of them topped the leaderboard. Five much smaller, purpose-built models did, and some of them have fewer than a billion parameters, a fraction of the size of the systems they beat.

That result comes from a paper in Cell, and it matters for reasons well beyond bragging rights. It is the first open, peer-reviewed attempt to test whether AI systems actually reason about ageing biology, rather than simply repeating facts about it that they memorised during training. For a field awash in AI-powered longevity claims, that distinction is overdue.

The study, "An open benchmark and language models for AI in aging biology," was the cover feature of Cell's 17 September issue. Insilico Medicine led the work with researchers from Liquid AI, the Buck Institute for Research on Aging, Harvard Medical School and Brigham and Women's Hospital, and the paper was published open access (Insilico Medicine, 2026; Zhavoronkov et al., Cell, 2026).

The team built three things and released all of them publicly. The first is LongevityBench: 17 tasks spanning five types of biological data (clinical records, genetics, epigenetics, gene expression and proteins), built so a model cannot simply recall an answer it saw during training. The second is a family of five Longevity-LLMs, compact language models fine-tuned specifically on ageing data. The third is Longevity Claw, an open-source research agent that uses those models to search for new drug targets on its own.

The headline result is really two results stacked together. When the team tested 18 frontier systems on LongevityBench, no single model dominated every task, and predicting age from omics data (patterns across DNA, RNA or proteins) was the hardest task at any scale, according to the paper. The researchers then fine-tuned their own small models, ranging from 0.6 billion to 9 billion parameters, on the same kind of data. Those compact models matched or beat every frontier system on the public leaderboard, which now tracks 26 models in total (LongevityBench leaderboard). On DNA-methylation age prediction, the strongest Longevity-LLM reached a concordance score of 0.868 against 0.685 for the best frontier model; on proteomic age prediction, a smaller Longevity-LLM cut the error from 10.1 years down to 5.7 years.

Ageing clocks and the multi-omics puzzle#

Every cell in the body carries evidence of how old it really is, and that age does not always match the number on a birth certificate. Researchers estimate this "biological age" using ageing clocks: statistical models trained to guess age from a blood sample, a set of proteins, or a DNA methylation profile. The best-known example is the Horvath clock, built from methylation, a chemical tag that switches genes on or off without changing the underlying DNA sequence. It is one of several "epigenetic" marks that drift in predictable ways as we age. Newer clocks work from blood proteins (proteomics) or gene activity levels (transcriptomics) instead, and each data type tends to capture a slightly different slice of the ageing process.

Large language models, the technology behind ChatGPT, Gemini and Claude, are trained to predict text. That makes them fluent explainers of biology, but fluency is not the same as reasoning. A model can describe what a methylation clock measures in perfect detail and still fail to read an actual methylation dataset and estimate someone's age from it correctly. That gap, between talking about data and correctly interpreting it, is exactly what LongevityBench was designed to expose. Its authors avoided public trivia questions a model might have memorised, testing instead against held-out numerical data from sources such as the NHANES health survey, the GEO methylation database, GTEx tissue samples and Olink protein panels.

Why this matters#

For AI research generally, the result is a useful check on hype. General-purpose chatbots get pitched as all-purpose scientific assistants, but LongevityBench shows that a strong reputation on coding or maths leaderboards does not automatically carry over to reading a proteomics file correctly. That is worth knowing before anyone builds a clinical tool on a general model without first checking its performance on the specific data type involved.

For ageing research, the more immediate benefit is access. Small, specialised models are cheap enough to run on ordinary hardware, which matters for academic labs that cannot afford frontier-model bills for every analysis. Insilico has released the benchmark, the five models and their training data under open licences on Hugging Face and GitHub (Hugging Face collection; GitHub), so any lab can test its own methods against a shared standard instead of inventing its own yardstick, which has long been a problem in ageing-clock research. As Insilico's founder and co-CEO Alex Zhavoronkov put it in the announcement, the goal is "benchmarked, agentic systems that can evolve into personalized longevity assistants and longevity companions, ultimately helping people monitor and improve their healthspan" (Insilico Medicine, 2026).

Critical analysis#

The strongest part of this work is the benchmark itself. By drawing on held-out clinical and omics datasets rather than published trivia, LongevityBench makes it harder for a model to fake competence through memorisation, a known weak spot in earlier biomedical benchmarks. Longevity Claw adds a genuine test of usefulness beyond the leaderboard: deployed across the platform's 14 tracked hallmarks of ageing, it nominated 328 candidate genes, enriched up to 5.6-fold for genes already linked to ageing in independent studies (Insilico Medicine, 2026). One nominee, KDM1A, had separately been shown to extend lifespan in C. elegans when modulated, which is a real external check rather than a number the authors produced themselves.

The limitations are, a 5.6-fold enrichment against a reference set is encouraging, but that reference set was itself built from earlier research, with its own gaps and blind spots, so Longevity Claw's picks inherit them. The benchmark tests reasoning over biological data, not whether a nominated gene will work as a drug target in an actual patient, and that gap between computational promise and clinical proof has swallowed a great many earlier candidates in this field. The compact Longevity-LLMs were also fine-tuned specifically for these task types, so it remains an open question how well their advantage holds up on ageing problems the benchmark does not cover. Realistically, expect years rather than months before any Longevity Claw nomination reaches a clinical trial. Insilico's own rentosertib, discussed below, needed a full drug-development programme, from AI-assisted target identification through to a Phase IIa readout, before it produced any human ageing-related result at all.

Expert perspective#

This is not Insilico's first pass at bridging AI and ageing biology, and it is not the field's first benchmark either. A narrower effort called ComputAgeBench tested 13 published clocks against 66 methylation datasets back in 2024, but it covered only epigenetic clocks, not the five data types LongevityBench spans, and it did not test general-purpose language models at all (Kriukov et al., ACM SIGKDD, 2025). Insilico's own group had previewed pieces of this work already: a January 2026 preprint introduced an early version of the benchmark, and a March 2026 preprint described a single 14-billion-parameter prototype that beat the Horvath clock on epigenetic age prediction. Both are preprints and were not peer-reviewed at the time of posting (Zhavoronkov et al., bioRxiv, January 2026; Zhavoronkov et al., bioRxiv, March 2026). The Cell paper is the peer-reviewed, scaled-up version of that trajectory: five models instead of one, a public leaderboard instead of an internal ranking, and an agent that acts on the results rather than just producing them.

It also arrives ten days after Insilico's other September headline: a Phase IIa trial in which rentosertib, an AI-designed drug for lung fibrosis, appeared to reduce biological age across six separate proteomic clocks in a subset of 42 trial participants (Insilico Medicine / Nature Biotechnology, 2026). Seen together, the two papers show a company backing both ends of the pipeline at once: a clinical asset moving through trials, and open infrastructure meant to help everyone else move faster too. Whether that infrastructure becomes a genuine field standard, in the way ImageNet once did for computer vision, will depend on whether outside labs actually adopt LongevityBench rather than building their own.

Key takeaways#

  • Cell published the first open, peer-reviewed benchmark testing whether AI systems can reason about ageing biology, not just describe it.
  • Eighteen frontier AI systems from six developers were tested. None dominated, and predicting age from omics data was the hardest task for all of them.
  • Five small, specialised Longevity-LLMs, from 0.6 to 9 billion parameters, matched or beat every frontier model once fine-tuned on ageing-specific data.
  • The accompanying Longevity Claw agent nominated 328 candidate ageing-intervention genes, with meaningful overlap against independently validated targets.
  • The benchmark, the models and the agent are all open source, giving labs without frontier-scale budgets a shared, harder-to-game way to measure progress.

Frequently asked questions#

What is LongevityBench? An open set of 17 tasks that tests whether AI systems can interpret clinical, genetic, epigenetic, transcriptomic and proteomic data well enough to draw correct conclusions about ageing, rather than just talking about the topic fluently (Zhavoronkov et al., Cell, 2026).

What are Longevity-LLMs? Five compact language models, between 0.6 and 9 billion parameters, fine-tuned on ageing-specific data and built on Liquid AI's LFM2 architecture and Alibaba's Qwen3 model family (Insilico Medicine, 2026).

What is Longevity Claw? An open-source AI agent that pairs a Longevity-LLM with tools for calculating biological age and screening candidate drug targets, running multi-step research workflows rather than answering one question at a time.

Why did small models beat much larger ones? Frontier models are trained mostly on general text and code. The compact models were fine-tuned specifically on the structured biological data the benchmark tests, and that specialisation seems to matter more than raw parameter count for this kind of task.

Does this mean AI has solved ageing? No. The benchmark measures how well a model interprets existing biological data. It does not show that any nominated drug target will work in humans, which still requires years of laboratory and clinical testing.

Is this connected to Insilico's rentosertib drug news from earlier in September? It is related but separate. Rentosertib's result came from a Phase IIa clinical trial published in Nature Biotechnology. The Cell paper is about open research infrastructure, not a drug trial, though both come from the same company within the same month.

Can other researchers use these tools? Yes. The benchmark, the models, the training data and the Longevity Claw code are all freely available on Hugging Face and GitHub under open licences.

Which AI systems were tested, and how did they rank? The paper tested 18 frontier systems from OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI. On the public leaderboard's aggregate score, Gemini 3.1 Pro ranks as the strongest frontier system, ahead of Claude Opus-4.6, though both trail the specialised Longevity-LLMs.

Glossary#

Biological age: An estimate of how old the body's tissues appear based on molecular or physiological measurements, which can differ from chronological age.

Ageing clock: A statistical or AI model trained to estimate biological age from data such as DNA methylation, blood proteins or gene activity.

Epigenetics: Chemical modifications, such as DNA methylation, that switch genes on or off without altering the underlying DNA sequence, and that shift in predictable ways with age.

Proteomics / transcriptomics: The large-scale study of all proteins (proteomics) or all active genes (transcriptomics) in a sample, used to build more detailed ageing clocks than a single measurement allows.

Foundation model: A large AI model, typically a language model, trained on broad data and then adapted for specific tasks.

Fine-tuning: Further training a general AI model on a narrower dataset so it performs better on a specific task, such as reading methylation data.

Agentic AI: An AI system that can plan and carry out several linked steps toward a goal on its own, rather than answering one question at a time.

Hallmarks of ageing: A set of biological processes, such as DNA damage accumulation and cellular senescence, that researchers use to categorise how and why the body ages.

References#

  1. Zhavoronkov, A. et al. "An open benchmark and language models for AI in aging biology." Cell, vol. 189, issue 19, pp. 5980-5994.e8, 17 September 2026. DOI: 10.1016/j.cell.2026.08.026. Peer-reviewed.
  2. Insilico Medicine. "Insilico Medicine Opens AI Longevity Discovery Toolkit to Researchers Worldwide in Cell Cover Study." PR Newswire, 17 September 2026. Link. Official organisation source.
  3. LongevityBench public leaderboard. longevitybenchmarks.org. Official project resource, updated on a rolling basis.
  4. Insilico Medicine / Liquid AI. Longevity-LLM and Longevity Claw model and code releases. Hugging Face collection; GitHub repository. Official/open-source resource.
  5. "Integration of proteomic aging clocks in a phase 2a clinical trial supports simultaneous geroprotective assessment." Nature Biotechnology, 7 September 2026. Link. Peer-reviewed.
  6. Kriukov, D., Efimov, E., Kuzmina, E. A., Khrameeva, E. E., Dylov, D. V. "ComputAgeBench: Epigenetic Aging Clocks Benchmark." Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2025. DOI: 10.1145/3711896.3737382. Peer-reviewed conference paper.
  7. Zhavoronkov, A. et al. "LongevityBench: Are SotA LLMs ready for aging research?" bioRxiv, posted 13 January 2026. DOI: 10.64898/2026.01.12.698650. Preprint (not peer-reviewed).
  8. Zhavoronkov, A. et al. "The End of Aging Clocks: Training Foundation Models to Reason in Aging and Longevity." bioRxiv, posted 30 March 2026. DOI: 10.64898/2026.03.28.714980. Preprint (not peer-reviewed).
  9. Bloom, A. "Insilico Medicine Releases Open Longevity AI Toolkit in Cell Study." Unite.AI, 17 September 2026. Link. Science journalism, used for supplementary detail only.

Related observations

Adjacent work from the same lines of enquiry.