Abstract
<title>Abstract</title> <p> <bold>Introduction.</bold> Depressive disorders are among the leading contributors to years lived with disability, and a large share of affected adults are neither identified nor treated. Machine-learning analyses of national survey data commonly rely on a single survey cycle, disregard the complex sampling design, and admit predictors that are consequences of depression rather than antecedents. <bold>Methods.</bold> Six two-year cycles of the National Health and Nutrition Examination Survey covering 2007–2018 were pooled. The outcome was a Patient Health Questionnaire-9 total of 10 or greater. One hundred and thirty-five candidate predictors were assigned in advance to nine domains, and item-level missing data were multiply imputed. Associations were estimated by survey-weighted logistic regression; six algorithms were compared on a held-out partition, with leave-one-domain-out ablation and a tiered analysis by data-collection cost. <bold>Results.</bold> Of 59,842 pooled participants, 31,460 adults formed the analytic sample and design-weighted prevalence was 8.05% (95% CI 7.58–8.54). A 67-item questionnaire matched the full 135-variable protocol (ROC-AUC 0.8454 against 0.8464). Four of nine domains carried unique predictive information. Penalised logistic regression outperformed all five machine-learning comparators. The strongest independently measured associations were reported trouble sleeping (OR 4.03), very low family income (OR 3.50) and food insecurity (OR 2.81). <bold>Discussion.</bold> Socioeconomic disadvantage, chronic disease burden and sleep disturbance are the most defensible correlates, and biological measurement added nothing detectable. <bold>Significance.</bold> Population screening for depressive symptoms does not require laboratory measurement, which carries a direct implication for screening cost. <bold>Future work.</bold> Longitudinal measurement is required to resolve the direction of these associations. </p>