Abstract
<sec> <title>BACKGROUND</title> <p>Smartwatches are increasingly used for decentralized data collection in clinical research, but the everyday-life settings that make these data attractive also introduce variability. Before a smartwatch-derived feature can support clinical monitoring, its reproducibility must be established: features with low test-retest reliability weaken associations with clinical outcomes and potentially generate non-actionable signals. Reliability is expected to vary by feature type, aggregation window, and data availability, but has not been systematically screened in a clinical cohort.</p> </sec> <sec> <title>OBJECTIVE</title> <p>This study aimed to evaluate the test-retest reliability of candidate smartwatch-derived features for longitudinal monitoring in adults with advanced cancer and in healthy controls, and to determine how reliability depends on the temporal aggregation window.</p> </sec> <sec> <title>METHODS</title> <p>In a prospective single-centre observational cohort study, we analysed 8 weeks of Garmin Vivosmart 5 sensor data from 60 adults with advanced cancer, and 20 healthy controls. Out of 80 participants, 77 contributed usable smartwatch data. We examined 35 daily features across seven domains: heart rate variability, heart rate, respiration, oxygen saturation, sleep, activity, and smartwatch-derived stress. Test-retest reliability was quantified as the intraclass correlation coefficient (ICC(2,1)) across adjacent non-overlapping 1-, 3-, and 7-day windows, with 95% CIs from subject-level bootstrap resampling (10,000 resamples). Between-group and therapy-centred contrasts used permutation testing with Benjamini-Hochberg false discovery rate correction.</p> </sec> <sec> <title>RESULTS</title> <p>Reliability improved with longer aggregation windows in both cohorts. Between 1-day and 7-day windows, median ICC(2,1) increased from 0.53 to 0.73 in controls and from 0.66 to 0.79 in patients. The number of features reaching good-to-excellent reliability, defined as ICC(2,1)≥0.75, increased from 5 of 35 to 15 of 34 in controls and from 12 of 35 to 26 of 35 in patients. Heart rate and heart rate variability features were the most reliable, with 4 of 11 reaching weekly ICC(2,1)≥0.90 in both cohorts. Activity, respiration, sleep and oxygen-saturation features were more sensitive to aggregation window and data availability, showing larger gains from daily to weekly aggregation (e.g., step count ICC increased from 0.30 to 0.73 in controls). Weekly reliability did not differ significantly between patients and controls (median ΔICC=0.036, no feature survived FDR correction, all q>0.05). No feature showed a significant change in reliability around therapy (median ΔICC=0.016, all q>0.05).</p> </sec> <sec> <title>CONCLUSIONS</title> <p>Weekly aggregation improved the reliability of many smartwatch-derived features, but reliability remained feature specific. Heart rate and heart rate-variability features were consistently reliable, whereas selected sleep and oxygen-saturation features displayed only moderate reliability across all aggregation windows. Reliability was comparable across cohorts and stable around therapy, indicating that feature-wise estimates are transferable across these clinical contexts. Feature-level reliability screening is a prerequisite before smartwatch-derived measures are used in clinical monitoring.</p> </sec>