Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Vision-Language-Action (VLA) foundation models have achieved significant progress in robotic manipulation. However, existing approaches suffer from a fundamental training-inference-execution disconnect. VLA-JEPA optimizes representation learning at pre-training but lacks online correction capacity, while VLA-Corrector performs inference-time error correction but cannot distinguish deviation types or provide semantic guidance for replanning. We present a Semantics-Aware Three-Layer Co-Optimization Architecture that bridges this gap through: (1) a Semantic Dynamics Monitor (SDM) that detects execution deviations in latent space and classifies them into three categories—transient noise, persistent execution drift, and structural environmental changes—based on temporal features; (2) a causal attribution module that identifies whether deviation originates from perception, dynamics modeling, or control; and (3) a Semantics-Anchored Replanner (SAR) that introduces language-instruction semantic consistency loss during gradient-guided replanning. Unlike prior one-size-fits-all correction mechanisms, our architecture enables deviation-type-adaptive intervention. Experiments on dynamic manipulation benchmarks demonstrate that our approach improves task success rates by 18.7 percentage points over VLA-Corrector under severe disturbances, while preserving native VLA performance in disturbance-free scenarios. Our framework is plug-and-play, requiring no retraining of the VLA backbone.</p>

Show More

Keywords

correction semantic manipulation from while

Related Articles

PORE

About

Connect