Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Safety monitoring during training is useful only if an internal warning can be shown to precede an equally sensitive output test. To our knowledge, existing studies separately analyze backdoor training dynamics, completed-model representations, or internal progress measures for benign capabilities; none establishes this causal temporal comparison. We present an executable protocol for doing so. The protocol adapts GPT-2 to SST-2 and then continued with a 3.125% activemarker poison or a matched-null condition that has identical inputs, marker exposure, label totals, and label noise. Every update is evaluated with (i) a continuous output-logit margin and (ii) a baseline-corrected post-block-10 residual projection onto a direction learned in an independent discovery continuation. Both tests use the same 128 paired prompts, a studentized mean, and one maximum-statistic cutoff over 160 detector-checkpoints. Causal status is tested by removing the acquired coordinate and restoring it with a checkpoint-matched amplitude from independent donors on a disjoint prompt set. Design analysis shows why this matching matters: using an unadjusted 5% test at 160 looks gives a 99.97% family-wise false-positive probability. A finite-sample Bonferroni boundary is t127 = 3.508, corresponding to an 80% minimum detectable standardized shift of 0.386 at n = 128 for either detector. A smoke test verified the primary training, evaluation, and final-checkpoint intervention path; synthetic unit tests verified the temporal causal-gate analysis. Both were expressly non-inferential. The substantive question therefore remains open; the contribution is a falsifiable criterion for answering it without mistaking token identity, thresholded behavior, probe flexibility, or unequal error budgets for an internal precursor.</p>

Show More

Keywords

training internal test causal temporal

Related Articles

PORE

About

Connect