Abstract
<jats:p>Proprioception can be assessed in several ways, including movement detection, joint position matching, and matching across sensory frames of reference. These task types make different demands, yet they are rarely compared within the same participants on the same device, and psychometric data for wrist-focused batteries are limited. This work had two aims: to compare performance across different levels of proprioceptive judgment, and to establish the within-day test-retest reliability of each. We evaluated three robotic wrist tasks spanning judgments within a single reference frame and across reference frames: joint detection threshold (JDT), same-frame joint-to-joint matching (J-to-J), and cross-frame joint-to-visual matching (J-to-V). Methods. Twenty neurotypical adults completed two identical sessions on the same day, separated by at least two hours, using a single-degree-of-freedom wrist robot. Outcomes were the kinematic detection threshold (degrees) for JDT and the mean absolute matching error (degrees) for J-to-J and J-to-V. Relative reliability was quantified with ICC (2,1) and 95% confidence intervals. Absolute reliability was quantified with the standard error of measurement (SEM) and the smallest detectable change at 95% confidence (SDC 95). Learning effects and differences across task levels were evaluated with paired t-tests or Wilcoxon signed-rank tests. Results. ICC (2,1) was 0.959 [95% CI: 0.900 to 0.980] for JDT, 0.837 [0.640 to 0.930] for J-to-J, and 0.769 [0.500 to 0.900] for J-to-V. The %SEM ranged from 11.9% (J-to-J) to 15.6% (J-to-V). SDC95 was 0.85, 1.71, and 3.54 degrees for JDT, J-to-J, and J-to-V, respectively. A small but significant practice effect was observed for JDT, but this was below the SDC95, and no learning effect was observed for J-to-J or J-to-V. We also observed that the absolute error increased monotonically across task levels, with all pairwise comparisons (JDT < J-to-J < J-to-V; all p < 0.01). Conclusions. All three tasks demonstrated good-to-excellent within-day relative reliability. Error scaled with the computational demand of each task, with the largest errors observed for the cross-frame task, which required a transformation between the visual and joint reference frames. The reported SDC95 values provide task-specific thresholds for distinguishing measurement noise from true change in future intervention studies. Inter-day reliability and validation in clinical populations are the next steps.</jats:p>