Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Monocular visual odometry (VO) estimates translation only up to scale. We compare five non-ground-truth scale priors while holding the SWiFT-VO geometry, a Depth Anything V2 relative-depth branch, and Qwen2.5-VL-7B fixed. A calibrated scene-depth heuristic transfers poorly, and three triangulation variants become unstable when the model distance and triangulated geometry refer to different surfaces. C3 avoids this mismatch by tracking one semantic anchor across keyframes and combining its VLM-reported distance with unit-baseline triangulation of the same anchor. Across eight KITTI sequences, C3 reaches a mean cumulative-distance ratio of 0.64. A fixed disparity gate rejects weak or implausible tracks and falls back to the heuristic, raising the mean ratio to 0.71; an online gate reaches 0.69. On the full KITTI 08 sequence, none of the VLM priors improves trajectory shape over GT-derived scale. The short-window results also show that ATE can favour under-scaling when the shared translation direction is biased. The study therefore reports partial metric-scale recovery and the conditions under which the anchor-based estimate is reliable.</p>

Show More

Keywords

scale translation priors geometry fixed

Related Articles

PORE

About

Connect