Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Background Diabetes affects tens of millions of Americans [1], and its prevalence varies markedly between neighboring communities; social determinants of health explain much of that variation independent of individual clinical risk [2,3]. Most predictive modeling in this space still relies on individual-level clinical data; community-level prediction using only public data remains comparatively rare, and few studies separate how much of a model's accuracy comes from actionable community characteristics versus from comorbidities that are themselves close correlates of diabetes. Methods We linked American Community Survey (ACS) socioeconomic indicators, USDA Food Access Research Atlas measures [10], OpenStreetMap built-environment densities, and CDC PLACES health estimates [14] across 412 census tracts in Collin and Denton Counties, Texas. Two feature sets were compared: a primary, actionable Model A (community, food-access, and built-environment variables only) and a Model B comparison that adds obesity, physical inactivity, and hypertension prevalence. Baseline, ordinary least squares (OLS), LASSO [17], and random forest [16] models were evaluated with an 80/20 split, 5-fold out-of-fold cross-validation, nested cross-validation with inner hyperparameter tuning [26], and geographic external validation (train Collin, test Denton). SHAP values [18,19] and OLS coefficients provided model interpretation; equity was assessed across income terciles; residual spatial structure was assessed with a Moran scatterplot [22]. Results Mean tract-level diabetes prevalence was 9.4% (SD 1.9; range 3.6–17.4%). Model A explained a modest, consistent share of tract-to-tract variation: R² = 0.519 (95% bootstrap CI 0.402–0.624) by cross-validation, 0.523 by nested cross-validation, and 0.592 under geographic external validation. Adding comorbidities (Model B) raised out-of-fold R² to 0.935 (95% CI 0.908–0.953), but SHAP attributed most of that gain to hypertension (r = 0.87 with diabetes prevalence), so we treat Model B as a comorbidity ceiling rather than the study's central finding. Model A's error was roughly twice as large in low-income tracts (MAE = 1.29) as in high-income tracts (MAE = 0.63; gap 95% CI 0.44–0.89, permutation p = 0.0001). Conclusions Public community, food-access, and built-environment data predict roughly half of tract-level diabetes variation, comparable to several recently published community- and county-level models [4,5,6], with reasonable stability out of county. Predictive accuracy is not equity-neutral, and that gap should temper any prioritization use of the resulting model.</p>

Show More

Keywords

model diabetes community prevalence crossvalidation

Related Articles

PORE

About

Connect