Abstract
<title>Abstract</title> <p>Breast cancer screening protocols apply uniform strategies across populations with vastly different underlying risk profiles, limiting early detection in high-risk individuals and generating unnecessary burden in low-risk ones. Risk-based screening has been proposed as a strategy to tailor screening intensity to individual risk, but its implementation requires accurate, calibrated, and scalable risk models validated in real-world settings. Here we develop and validate a population-scale breast cancer risk model using longitudinal mammography reports and clinical data from approximately 4.6 million examinations of 1.8 million patients across the largest private health system in Brazil. Without requiring mammography images, the model achieves strong and clinically meaningful discrimination for predicting breast cancer within one to five years, with a 5-year AUROC of 0.86 (95% CI: 0.85–0.87) and consistent performance across shorter horizons. Flagging the top 3.4% highest-risk exams identifies 50.1% of future cancer cases, enabling earlier detection, while in the bottom 47% risk percentile only 10.0% develop cancer within five years, suggesting approximately half of all patients could safely transition to longer screening intervals. An explainability analysis reveals distinct and clinically meaningful patterns driving predictions at both ends of the risk spectrum. These results establish the feasibility of interpretable, image-free, population-scale risk modeling as a practical pathway toward personalized breast cancer screening.</p>