Abstract
<jats:p>Abstract. High-resolution convection-permitting models (CPMs) provide vital datasets for studying localized extreme precipitation and associated multi-hazard risks like landslides and floods. Yet validating CPM datasets against observations remains a challenge. This study evaluates two sub-daily CPM datasets, VHR-REA_IT and LAM-HIND, against observations over a 10-year period across two Italian regions: Emilia-Romagna and Campania. Extreme precipitation is defined using the 95th percentile of wet-event intensities across accumulation durations of 1, 2, 4, 12, and 24 hours. Conventional verification metrics only partially assess dataset performance and hide differences in the temporal organization of precipitation extremes. Therefore, a new nested time-window framework is presented to better capture the association between short- and long-duration time window extremes (STWEs and LTWEs, respectively). Observations show that STWEs and LTWEs are only partially associated. Depending on region and season, around 7–15 % of observed 1-hour STWEs are embedded within 24-hour LTWEs. CPMs reproduce this cross-timescale organisation substantially better than ERA5, yet are overestimated, indicating excess temporal persistence of simulated heavy-rainfall. Consequently, LTWEs are more dependable than STWEs, as CPMs distribute short extreme precipitation bursts into longer durations with persistently higher volumes. Differences between the CPM datasets are small. LAM-HIND exhibits subtly higher detection and event-overlap scores but a higher false-alarm rate and stronger cross-timescale coupling overestimation. VHR-REA_IT more closely reproduces the observed temporal association between short- and long-duration extremes. These differences may stem from the ‘double penalty’ effect, choice of an intermediate nesting strategy, and sensitivity in resolving sub-grid turbulent fluxes. However, they are both substantially improved over coarse products like ERA5 for sub-daily extreme precipitation.</jats:p>