Abstract
<title>Abstract</title> <p>Safety alignment and helpfulness training are usually treated as objectives to be combined, but their order may matter because gradient updates in a neural network need not commute. We isolate this effect in a small open instruction model by holding the training examples, objectivespecific minibatches, update counts, and initial parameters fixed. We compare three schedules: safety direct preference optimization (DPO) followed by helpfulness supervised fine-tuning (SFT), the reverse order, and an alternating control. Across dense checkpoints, we measure refusal of unsafe requests, over-refusal of benign requests, generalization to paraphrases, and held-out helpfulness. We then test whether the behavioral change is associated with a stable representation distributed through the residual stream or with a change concentrated near the output layers. Ending with helpfulness rather than safety reduced unsafe-prompt refusal by 65.8 percentage points (95% hierarchical bootstrap interval: 60.1–71.0 points) and benign over-refusal by 36.8 points, while improving held-out completion NLL by 0.028 nats/token. The interleaved schedule was between the two block orders. Raw harm labels remained linearly decodable, but the training-induced direction rotated in upper layers. A safety-anchor direction did not transfer as a stable causal mediator; an endpoint-specific projection swap produced a modest effect near block 28 after safety-last and interleaved training, but not after helpfulness-last training. The experiment provides a matched test of path dependence in post-training and a template for connecting alignment behavior to mechanisms as they develop.</p>