An open benchmark and language models for AI in aging biology
LongevityBench is the first open benchmark testing whether AI systems can actually interpret aging biodata, spanning 17 tasks across five domains and 18 frontier models from six developers. No model dominates, and omics-based age prediction is the hardest task regardless of model scale - directly relevant given how much clinical longevity practice now routes through algorithmic age estimates. Compact fine-tuned Longevity-LLMs (0.6B-9B parameters) matched or exceeded much larger systems. Infrastructure, not a clinical claim; note the commercial affiliation.
Evidence
5/10
Emerging Evidence
Sample
—
subjects
Duration
—
study period
Journal
Cell
Sep 2026
Full Abstract
Over the past two decades, human aging has been characterized across DNA methylation, transcriptomic, proteomic and clinical modalities, yet no benchmark evaluates whether AI systems can interpret these heterogeneous data types in the context of aging biology. The authors introduce LongevityBench, an open suite of 17 tasks spanning five biodata domains, and use it to assess 18 frontier AI systems from six developer teams. No single model dominates all tasks, with omics-based age prediction being the hardest task regardless of scale. Fine-tuned compact multitask Longevity-LLMs (0.6B-9B parameters) matched or exceeded far larger frontier systems. The benchmark, models and an agentic research interface are publicly released.
Key Findings
- 01
LongevityBench: 17 open tasks spanning five aging biodata domains
- 02
18 frontier AI systems from six developer teams evaluated
- 03
No single model dominates across tasks
- 04
Omics-based age prediction is the hardest task irrespective of model scale
- 05
Compact fine-tuned Longevity-LLMs (0.6B-9B parameters) matched or exceeded far larger frontier systems
- 06
Benchmark, models and an agentic research interface released publicly
Structured Methods
- Study Design
- Diagnostic Study
- Sample Size
- Not reported
- Study Duration
- Not reported
- Methodology
- Construction of a 17-task open benchmark across DNA methylation, transcriptomic, proteomic, clinical and literature domains; evaluation of 18 frontier AI systems from six developer teams; fine-tuning of five compact multitask domain-specific language models for comparison.
- Limitations
- Produced by Insilico Medicine, a commercial longevity-AI company benchmarking a field in which it competes. A benchmark's task selection is itself a scientific choice and may favour certain architectures. Measures AI interpretive capability, not clinical validity of any aging biomarker, and carries no patient-level implication.
Citations & References
Alex Zhavoronkov, Vladimir Naumov, Diana Sidorenko, Alex Aliper, Vladimir Aladinskiy (2026). An open benchmark and language models for AI in aging biology. Cell. https://doi.org/10.1016/j.cell.2026.08.026
Sample member
Longevity Operator
One number for your longevity journey.
FICO for credit. PHS for longevity. Eight biological domains collapsed into one provable score — intuitive at a glance, rigorous underneath.
- Metabolic
- Hormonal
- Cardiovascular
- Brain / Cognitive
- Sleep
- Musculoskeletal
- Gut
- Immunity / Inflammation
