Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Deep learning is now routinely used to build biomarkers of biological age, and the field has largely equated methodological progress with increases in model capacity—from penalized linear clocks to deep neural networks and, most recently, foundation models. Yet a recurring observation is that added capacity buys accuracy at predicting chronological age without necessarily buying a better biomarker. We test one narrow, falsifiable version of this tension on routine blood chemistry. Using the National Health and Nutrition Examination Survey (NHANES III as training, n=9,200; Continuous NHANES 1999–2010 as external validation, n=16,400) with National Death Index–linked all-cause mortality (median follow-up 20.6 years; 4,368 validation deaths), we build six aging clocks in a fully crossed 2 × 3 design: three model capacities (elastic net, gradient-boosted trees, a deep multilayer perceptron) crossed with two prediction targets (chronological age versus a mortality-derived phenotypic age). We evaluate every clock on three axes—chronological-age prediction accuracy, test–retest reliability, and biomarker validity, defined as the mortality association of the age gap (age acceleration) beyond chronological age—and decompose the between-clock variance in validity into capacity and target components. The results dissociate sharply. Higher-capacity clocks predict chronological age slightly better in internal cross-validation (deep MLP MAE 5.06 yr versus elastic net 5.43 yr), but the advantage nearly vanishes out of sample (both ≈ 5.48 yr) and higher capacity lowers test–retest reliability (ICC 0.970 versus 0.985). Biomarker validity is governed almost entirely by the target: the age-gap hazard ratio per standard deviation, adjusted for chronological age and sex, is 1.14–1.23 for chronological-age clocks and 1.48–1.59 for phenotypic clocks, and switching target multiplies the hazard ratio by 1.29× (95% CI 1.26–1.33) whereas moving from linear to deep multiplies it by 0.93× (i.e., no improvement). A two-way variance decomposition attributes 95% of the betweenclock variance in age-gap log-hazard-ratio to the target and 5% to capacity. The dissociation is stable across internal and external validation and within sex strata. For blood-based aging clocks within this regime, the prediction target—not model capacity—determines biomarker validity, and we argue this reframes where deep-learning effort should be spent.</p>

Show More

Keywords

clocks deep chronological capacity biomarker

Related Articles

PORE

About

Connect