Abstract
<title>Abstract</title> <p>Artificial cognitive systems may need to infer how a demonstrated state was produced and apply that process to a new scene. We introduce a language-free task in which paired demonstrations share identical initial and final images but differ in one intermediate image D1, which identifies whether recoloring preceded movement or vice versa. The same query has two process-dependent outcomes. To separate learned intervention rules from evaluationtime evidence, we crossed paired-D1 substitution supervision and masked-D1 uncertainty supervision in a 2 × 2 ablation. The confirmatory study used 12 independently generated corpora and two paired initializations per corpus, with inference over corpus means. A jointly tuned explicit two-state system achieved .981 foreground intersection over union (IoU) and .927 exact-image accuracy on held-out positional parity, exceeding an endpoint-only control by .384 IoU (95% corpus-bootstrap interval [.376, .391]). The canonical-only regime received neither intervention loss, yet substituting the paired intermediate image reversed the explicit process and selected the paired alternative outcome with probability .99988 [.99968, .99999]. Substitution supervision added no measurable benefit. Masked uncertainty did not emerge from canonical training: absolute deviation from .5 was .492 without mask supervision and .001 with it (factorial effect −.493 [−.496, −.489]). True, soft, and hard process codes produced identical item-level discrete images. A high-bandwidth latent interface remained slightly stronger than the explicit interface (IoU difference −.008 [−.015, −.001]). These results demonstrate intervention-sensitive visual process control beyond explicitly trained substitution rules while distinguishing substitution consistency from supervision-dependent uncertainty calibration</p>