Abstract
<title>Abstract</title> <p>Executable world-model agents turn interaction histories into executable models that can be tested and searched, but predictive agreement does not establish causal correctness or planning adequacy. We introduce THEA (Typed Hypothesis-driven Epistemic Agents), a prompt-defined revision policy that treats the executable world model as a falsifiable task theory. A parent agent preserves causal context and retains exclusive authority over environment actions. Observable failures and declared checkpoints activate bounded specialists (subagents) for framing, generalization, raw-evidence review, rival-hypothesis testing, causal reinduction, and transition observation. Typed artifacts record competing hypotheses, discriminating probes, and verdicts, while the executable theory, logs, and interaction archive preserve the evidence behind each revision. We evaluate THEA, with a Claude Code Opus 4.8 parent at xhigh effort and provider-native Sonnet 5 specialists, on the 25 public ARC-AGI-3 games through a historical portfolio recording each game's final trajectory from across the harness's development. It fully solves 22/25 games, averages 92.63 per-game Relative Human Action Efficiency (RHAE), and has an API-equivalent cost of $5,007.85 versus $6,222.18 for the strongest published agent on this benchmark, the version 1.5 Executable World Models release (hereafter v1.5), a 19.52% reduction concentrated in games that v1.5 leaves incomplete. Archived cases show exact replay preserving a wrong causal proxy, exhaustive search closing under an incomplete model, a faithful engine feeding an invalid planner, and a correct theory arriving after an irreversible decision. THEA makes hypothesis revision explicit, falsifiable, and auditable while exposing where that revision succeeds, fails, and incurs cost, offering a control layer for complex, orchestrated agent systems. The harness, evidence archive, and per-game run records are available at https://github.com/bruno353/thea.</p>