Abstract
<title>Abstract</title> <p>Mechanistic interpretability usually asks what a trained model does at its final checkpoint. This leaves unused a second source of evidence: how the mechanism came together. I test whether multiple circuit states improve a language-model analyst’s structured mechanism recovery when the final evidence, task, model, and harness settings are held fixed. The benchmark contains 24 blinded, constructed causal dossiers from four graph families. Each dossier has six opaque components, an exact directed graph, forced-choice component roles, a five-state assembly history, and eight held-out post-intervention outputs. In the multiple-state condition, a GPT-5.6-Terra analyst received five ordered state panels. In the control condition, it received five independent measurement panels from the final state. The fifth panel was numerically identical across paired conditions, prompt lengths differed by two characters, and both conditions used the same model settings, timeout, output schema, and tool policy; no tool event occurred. Multiple-state evidence raised causal-edge average precision from 0.888 to 0.971 (paired difference 0.082, family-stratified 95% bootstrap interval [0.045, 0.122], Holm-adjusted p = 0.006) and held-out intervention score from 0.891 to 0.954 (difference 0.062 [0.033, 0.091], adjusted p = 0.003). Component-role macro-F1 was 1.000 in both conditions, creating a ceiling. In a post hoc eight-instance audit, masking the timestamps of intermediate states did not lower mean performance, but the interval was wide. Thus, for this model and benchmark, access to several constructed states improved edge ranking and post-intervention prediction; the experiment does not establish a benefit from chronological labels. The dossiers are causal abstractions motivated by compiled transformers, not optimization trajectories, so the result is a controlled proof of concept rather than evidence about natural language-model training.</p>