Abstract
<jats:p>Hi-C experiments produce genome-wide chromatin contact maps from which structural features can be quantified across multiple genomic scales, ranging from megabase-scale compartments to kilobase-scale boundaries and chromatin loops. Although high-resolution analyses commonly rely on hundreds of millions of sequenced read pairs, the minimum sequencing depth required for different classes of Hi-C-derived features has not been systematically established. We therefore sought to determine these depth requirements systematically. We performed progressive random subsampling of twelve Hi-C libraries representing multiple cell types and experimental protocols. Loop-density and loop-size inference from the P(s) log-derivative was evaluated across all libraries, whereas compartment, insulation, and boundary analyses were performed on a subset of nine libraries with comparable full-depth coverage. Rather than defining sufficient depth as an arbitrary fraction of the full-depth value, we introduced a biologically motivated criterion based on the variability between independent biological replicates. The required sequencing depth for each metric was defined as the point at which subsampling scatter first reached the variability observed between independent biological replicates, representing the accuracy that additional sequencing cannot improve upon. Using this criterion, loop-density estimation reached its biological floor at approximately 10 million read pairs, loop-size estimation at approximately 20 million, insulation scores and boundary detection at approximately 30 million, and compartment eigenvectors at 100-kilobase resolution at approximately 60 million read pairs. Different Hi-C-derived metrics require substantially different sequencing depths to achieve biologically meaningful accuracy. For all metrics, depth-induced variability fell below biological replicate variability well before full sequencing depth was reached. These thresholds provide practical guidance for experimental design and sequencing budget allocation, suggesting that, in many studies, increasing the number of biological replicates is likely to improve reproducibility more effectively than sequencing individual libraries to greater depth.</jats:p>