All Research
    PAPER LONGEVDiagnostic Study2026

    An open benchmark and language models for AI in aging biology

    LongevityBench is the first open benchmark testing whether AI systems can actually interpret aging biodata, spanning 17 tasks across five domains and 18 frontier models from six developers. No model dominates, and omics-based age prediction is the hardest task regardless of model scale - directly relevant given how much clinical longevity practice now routes through algorithmic age estimates. Compact fine-tuned Longevity-LLMs (0.6B-9B parameters) matched or exceeded much larger systems. Infrastructure, not a clinical claim; note the commercial affiliation.

    Evidence

    5/10

    Emerging Evidence

    Sample

    subjects

    Duration

    study period

    Journal

    Cell

    Sep 2026

    Authors

    Authorship

    Alex Zhavoronkov, Vladimir Naumov, Diana Sidorenko, Alex Aliper, Vladimir Aladinskiy

    01

    Full Abstract

    Over the past two decades, human aging has been characterized across DNA methylation, transcriptomic, proteomic and clinical modalities, yet no benchmark evaluates whether AI systems can interpret these heterogeneous data types in the context of aging biology. The authors introduce LongevityBench, an open suite of 17 tasks spanning five biodata domains, and use it to assess 18 frontier AI systems from six developer teams. No single model dominates all tasks, with omics-based age prediction being the hardest task regardless of scale. Fine-tuned compact multitask Longevity-LLMs (0.6B-9B parameters) matched or exceeded far larger frontier systems. The benchmark, models and an agentic research interface are publicly released.

    02

    Key Findings

    1. 01

      LongevityBench: 17 open tasks spanning five aging biodata domains

    2. 02

      18 frontier AI systems from six developer teams evaluated

    3. 03

      No single model dominates across tasks

    4. 04

      Omics-based age prediction is the hardest task irrespective of model scale

    5. 05

      Compact fine-tuned Longevity-LLMs (0.6B-9B parameters) matched or exceeded far larger frontier systems

    6. 06

      Benchmark, models and an agentic research interface released publicly

    03

    Structured Methods

    Study Design
    Diagnostic Study
    Sample Size
    Not reported
    Study Duration
    Not reported
    Methodology
    Construction of a 17-task open benchmark across DNA methylation, transcriptomic, proteomic, clinical and literature domains; evaluation of 18 frontier AI systems from six developer teams; fine-tuning of five compact multitask domain-specific language models for comparison.
    Limitations
    Produced by Insilico Medicine, a commercial longevity-AI company benchmarking a field in which it competes. A benchmark's task selection is itself a scientific choice and may favour certain architectures. Measures AI interpretive capability, not clinical validity of any aging biomarker, and carries no patient-level implication.
    04

    Citations & References

    Cite this paper

    Alex Zhavoronkov, Vladimir Naumov, Diana Sidorenko, Alex Aliper, Vladimir Aladinskiy (2026). An open benchmark and language models for AI in aging biology. Cell. https://doi.org/10.1016/j.cell.2026.08.026

    05

    Indexing

    Topics

    artificial intelligenceaging clocksbenchmarkingomics

    Interventions

    biological age testingepigenetic clocks
    PHS
    671
    / 1000
    T3
    Longevity Operator

    Sample member

    Longevity Operator

    The Peak Human Score · PHS v1

    One number for your longevity journey.

    FICO for credit. PHS for longevity. Eight biological domains collapsed into one provable score — intuitive at a glance, rigorous underneath.

    • Metabolic
    • Hormonal
    • Cardiovascular
    • Brain / Cognitive
    • Sleep
    • Musculoskeletal
    • Gut
    • Immunity / Inflammation