Abstract
<title>Abstract</title> <p>Purpose Mixed-reality (MR) surgical navigation requires an intraoperatively measured organ surface for deformable image-to-physical registration, yet commercial headsets often rely on manual stylus digitization or external sensors. This work evaluates whether Magic Leap 2 (ML2) onboard sensors can provide this surface without external hardware, and whether its depth sensor or an RGB learning-based reconstruction is more accurate. Methods Two paradigms were developed from identical ML2 input in a common SLAM-tracked frame. The first used the onboard short-range indirect time-of-flight (iToF) stream. The second recovered geometry from an RGB sweep using Depth Anything 3, with metric scale anchored by SLAM-tracked camera centers. Both were evaluated against optically tracked stylus ground truth on opaque plaster and translucent silicone liver phantoms, a surface-treated breast phantom, and an ex vivo porcine liver, with five runs per specimen. Sensor-based feasibility was also demonstrated in vivo on an anaesthetized pig. Results ML2 world-origin drift remained below 2 mm median across four perturbations, and the iToF sensor achieved approximately 2 mm absolute accuracy over 0.30–0.70 m. Accuracy depended on surface optics. The sensor-based method was more accurate on opaque plaster livers (RMSE 2.6 and 3.1 mm versus 4.7 and 5.0 mm), whereas the learning-based method was more accurate on translucent silicone livers (4.3 and 4.6 mm versus 4.8 and 5.3 mm). The sensor-based method also performed better on the sunscreen-coated breast phantom (2.9 versus 3.9 mm) and retained a modest advantage on the ex vivo porcine liver (4.2 versus 4.8 mm). In vivo, the sensor cloud showed a mean residual of 3.4 mm to its fitted surface. Conclusion A commercial MR headset can acquire intraoperative organ surfaces using onboard sensors alone. Sensor-based reconstruction is preferred when active depth returns are reliable, whereas RGB learning-based reconstruction is more robust to optical failure modes but requires adequate multi-view coverage.</p>