Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Accurate dynamics models are fundamental to model-based control of legged robots, yet existing neural network approaches suffer from autoregressive error accumulation over long horizons. Joint-Embedding Predictive Architectures (JEPAs) offer an alternative by learning dynamics in a compact latent space without reconstructing observations. Here we introduce Ground-JEPA, a minimal (~1.300,000 parameters) latent world model that learns to control legged robots entirely through latent consistency, without task-specific rewards. The key innovation is a physics-inspired probe that maps latent representations to interpretable physical states (position, velocity, orientation), enabling physically grounded prediction and interpretability. Challenging prior assumptions that reward-free learning degrades in high-dimensional systems, we show that <bold>extending the planning horizon to capture full gait cycles (H=12) and adopting a lightweight Transformer predictor not only closes but surpasses reward-grounded performance</bold> in high-DoF environments (quadruped-walk: 510 ± 68 vs. 450 ± 234, p &lt; 0.05). On humanoid-stand (21-DoF), we achieve 350 ± 50 the first demonstration of reward-free JEPA control on a humanoid. Under extreme compound dynamics shifts (mass ±30%, friction 2.5×, actuator latency 50ms), our model retains 78% of nominal performance zero-shot, establishing latent consistency as a robust self-supervisory signal for sim-to-real transfer. Our results overturn the assumption that reward-free learning is inherently inferior, providing a laptop-scale, interpretable pathway for advanced robot control that democratizes access to state-of-the-art model-based methods. </p>

Show More

Keywords

latent control dynamics learning rewardfree

Related Articles

PORE

About

Connect