Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>In-context learning (ICL) and in-weights learning (IWL) can solve the same prediction problem by using different information. Prior work shows that ICL can appear and later give way to a weight-based strategy, but this one-way change does not establish whether the current data distribution uniquely determines the mechanism. We test a stricter question: when a small transformer reaches the same ICL/IWL training mixture from opposite directions, does it show the same behavior and the same sensitivity to prespecified causal interventions? We construct an ambiguous associative-recall task in which the input distribution is fixed while a mixture parameter controls whether the target follows a contextual association or a stable association learned in the weights. From a shared checkpoint, paired models traverse counterbalanced triangular mixture schedules with identical update counts and identical multisets of training examples, but opposite final approach directions. Conflict trials measure the model’s continuous preference for contextual versus weight-based answers. Residual-stream activation patching, attention-head and MLP mean replacements, and fixed-mixture endpoint holds test mechanism use and relaxation. At the balanced mixture, arrival from below rather than above increased the immediate contextual-minus-fixed margin by 0.0976 (95% paired CI [0.0818, 0.1133]); all five seed contrasts had the same sign, giving the smallest attainable exact two-sided sign-flip 𝑝-value, 0.0625. The contrast briefly reversed after 25 common-mixture updates and was near zero after 100. Replacing all query-position attention-head outputs with checkpoint-specific means attenuated 86.8% of the held-out immediate branch gap. Yet paired models selected the same dominant head in every seed, their eight-head effect profiles correlated at 0.99974–0.99998, and their activation-patching profiles were nearly identical. Thus training order left a measurable but transient difference in the strength of a shared contextual route. We find no evidence for persistent quasi-static hysteresis or for different causal circuit organization at the matched mixture.</p>

Show More

Keywords

same mixture training from contextual

Related Articles

PORE

About

Connect