Back to Search View Original Cite This Article

Abstract

<jats:p>Background: Polygenic scores (PGSs) are increasingly used to investigate the genetic architecture of complex traits. In genetics, family study designs are often used to adjust for confounders such as population structure and shared environment. However, family studies may also be particularly vulnerable to to non-random ascertainment, for example when individual case status affects the probability of inclusion, leading to differential representation of sibling pairs. In sibling samples, PGS associations can be decomposed into within-family and between-family components, where the within-family estimate captures associations between sibling differences in PGS and differences in outcome, thereby providing an estimate that is less affected by shared familial confounding. In this study, we examined the impact of non-random sampling on estimated genetic effects in family-based studies using both simulations and real-world data. Further, we leveraged the iPSYCH study design to estimate within-family PGS effects for six common mental health outcomes and whether accounting for these can improve prediction accuracy. Methods: We conducted simulations and applied the same framework to real-world data to evaluate the impact of selection bias on within- and between-family PGS estimates. Selection bias was modelled through differential sampling of sibling pairs based on case status, and inverse probability weighting (IPW) was applied to adjust for known heterogeneous inclusion probabilities. Analyses were replicated in the iPSYCH cohort using registry-based sampling weights and PGSs for six major psychiatric disorders. Predictive performance of models was assessed using five-fold cross-validation. Results: In simulation studies, biased sampling led to deviations in estimated PGS effects, with greater distortion observed for between-family components. IPW adjustment reduced the discrepancy between estimates obtained from biased and true underlying data. In the iPSYCH cohort, between-family estimates from unweighted models were larger than within-family estimates across traits. IPW weighted attenuated several of these estimates. Prediction analyses comparing models using total PGS versus decomposed within- and between-family components showed minimal differences in area under the curve and scaled R2 in the iPSYCH data, while modest gains were observed in selected simulation scenarios. Conclusions: Non-random ascertainment distorts effect estimates in family-based models, with particular sensitivity when estimating between-family effects. Incorporating IPWs derived from known or estimable inclusion probabilities can reduce this bias. Our findings highlight the importance of accounting for selection bias in family studies when estimating genetic effects</jats:p>

Show More

Keywords

betweenfamily estimates effects studies sibling

Related Articles

PORE

About

Connect