Abstract
<jats:p>Robot-learning benchmarks that hide dynamics parameters—mass, friction, payload—have driven a mature line of adaptation methods: online system identification, meta-RL, and rapid motor adaptation. A distinct, software-level failure mode is equally real and structurally different. The mapping from a policy’s action tensor to controller behavior—which channel drives which axis (or mutation), each channel’s sign and scale, whether the target is a delta or an absolute pose, its reference frame, its pipeline lag, and the gripper convention—can itself be wrong, undeclared, or silently changed, and we are not aware of a prior benchmark that isolates this action-interface contract as the object of study while holding task dynamics fixed. We introduce ActionShift, a real-simulation benchmark on frozen, competent ManiSkill PPO and Diffusion Policy backbones (PickCube, PushCube, PullCube, StackCube) in which a discrete, compositional interface contract is hidden behind a leakage-safe wrapper across preregistered seen, unseen-composition, and long-lag splits, with a privileged oracle and preregistered promotion gates. A competent policy that is near-perfect under a known interface (up to 1.000) collapses to a 0.000–0.007 floor when it is hidden—identically for RL and imitation policies—and the same policy-agnostic belief adapters restore both. Across a matched-privilege tournament, entropy-guided probing clears our promotion gate over fixed probing on 2 of 4 tasks (+5.8 points, 3-seed paired-significant on PickCube), and our controller, DualABI, matches the champion’s success at roughly half the probe cost on all four tasks. Grammar-only belief, with no privileged pool, converts the learned-identifier tier’s ~0.0 into 0.36–0.67 and—via a hold-probe excitation schedule—breaches the all-absolute wall (0.005→0.722, 0.000→0.790). We report the load-bearing negatives in equal detail: passive learned identification fails outright (0.0), every reactive method including the oracle collapses under lag, and a delay-aware backbone solves the long-lag split by planning through delay (~20×/~2.7×)—evidence the two adaptation problems are not interchangeable in practice. We make no literature-priority (“first”) claims: every novelty statement is an explicit “we are not aware of” backed by a direct literature check.</jats:p>