Abstract
<title>Abstract</title> <p>A tool-using model does not need execution authority to create risk; a structured proposal is enough if the runtime dispatches it without a reliable check. This paper studies that check. PolicyFaultBench combines deterministic fail-closed mediation, state-transition oracles, proposal-interface checks, and mutation testing. Providers transcribe an operation rather than plan an open-ended action, so the benchmark measures runtime mediation and interface conformance, not general agent safety. The 40-case corpus exposed 7/12 ordered-rule mutants; five labelled survivor-guided probes exposed the remainder. Three OpenAI and Anthropic ledgers contributed 1,200 records, including a prospectively frozen 400/400 Anthropic confirmation under a corrected pre-policy gate. To test whether the findings depended on one constructed corpus or one policy architecture, a second extension was frozen before execution. It derived 40 synthetic operations from four AgentDojo v1 suites, added an implemented capability policy, and scheduled 400 trials for each of OpenAI and Anthropic. OpenAI met the compound endpoint in 400/400 trials. Anthropic met it in 391/400: nine outputs for one benign payment case added an unrequested approval-scope field and were quarantined before policy evaluation. Thus the preregistered 800/800 acceptance rule was not met, although all 791 admitted proposals agreed under both policies, every verified transition matched, and no containment failure occurred. Preserved core proposals also replayed 800/800 under the capability policy. External-corpus mutation scores were low before targeted probes, showing that cross-corpus execution success and policy fault adequacy are different properties. Runtime assurance therefore also needs separate evidence for interface exactness, policy correctness, state effects, and corpus-policy generalisation.</p>