Abstract
<title>Abstract</title> <p>Models that predict a person’s chronological age from routine blood chemistry are widely described as measuring “biological age”. The claim only means something if the error such a model makes — predicted minus actual age, usually called age acceleration — carries information about how long the person will live, over and above their actual age. We tested that claim directly. We joined the standard biochemistry and complete blood count panels from ten cycles of the National Health and Nutrition Examination Survey (1999–2018) to the public-use linked mortality file, giving 45,033 adults aged 20–79 with 5,213 deaths over 445,961 person-years. We refit chronological-age predictors from six model families under nested cross-validation, including a sweep of 360 neural networks, extracted age acceleration from each, and tested all of them in Cox models adjusted for chronological age and sex. Flexible models predict age much better than linear ones: out-of-fold R2 rose from 0.398 for ridge regression to 0.549 for gradient boosting and 0.557 for a neural network. The usefulness of their residuals moved the other way. Age acceleration from ridge regression raised Harrell’s Cindex for all-cause mortality by +0.0124 over age and sex, while age acceleration from the neural network raised it by only +0.0056. Across the six families the correlation between age-prediction R2 and C-index gain was −0.94 (p = 0.005): the more accurately a model hit chronological age, the less its residual was worth. Meanwhile the same 38 analytes fitted against mortality instead of against age gained +0.0367 (95% CI +0.0332, +0.0408) — 3.0 times the best age-trained residual — and retained 85% of that gain even after age acceleration was already in the model. Published PhenoAge acceleration, which uses only nine markers but was built against mortality, gained +0.0182. The limitation is therefore neither the biomarkers nor the model class; it is the training target. Finally, we show that the R2 values reported for blood-chemistry age predictors are mostly a property of the sample: holding the model and the analytes fixed and varying only the age window moves R2 from 0.247 to 0.683. Fitting a model to chronological age is a reasonable way to describe how blood chemistry changes with age; it is a poor way to build a mortality biomarker, and reporting R2 alone hides that.</p>