Abstract
<title>Abstract</title> <p> <bold>Purpose:</bold> Automated educational systems increasingly route rubric decisions on structured key fields, yet it is difficult to tell whether a model has learned the intended lookup or exploited surface cues. We ask whether a rubric-routing layer can be validated before deployment under a fully known synthetic mapping, and how the surface distance between the fields that must be combined affects that routing. <bold>Methods:</bold> We crossed four neutral source keys with four evidenceprofile keys to form 16 balanced associations to four action classes, and trained from-scratch character-level Transformers with a fixed key-span classifier. Across 144 frozen runs we compared three encodings of the same mapping—atomic keys, adjacent factorized keys, and factorized keys separated by 256 neutral spaces—at two model sizes, using prespecified noninferiority, distance, and intervention tests. <bold>Results:</bold> Atomic positive controls passed at both model sizes. In the larger model, adjacent factorization was noninferior to atomic routing, whereas the spacing manipulation reduced final accuracy by 0.250 and normalized learning-curve area by 0.0419. The smaller-model replication reproduced the spacing, timing, and intervention effects but not noninferiority. Source and profile substitutions shifted the classifier log odds toward the replacement-prescribed action. <bold>Conclusion:</bold> Surface key distance, rather than mapping content, governed routing in this controlled interface. The paradigm isolates a lookup mechanism rather than semantic assessment, and motivates synthetic unit tests as a diagnostic step before evaluation on authentic text and operational architectures. </p>