Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Scientific causal inference is frequently data limited: a target study may contain only tens of randomized or quasi-experimental observations, while related domains contain much larger observational and interventional datasets. We investigate whether intervention-rich cross-task pretraining can provide a reusable causal prior that improves treatment-effect estimation in a new small-data domain. We introduce SciCFM, an episodic causal pretraining framework that encodes support sets from heterogeneous structural causal models and adapts the learned mechanism representation to a target dose-response task. Three simulation studies were conducted. First, under changing hidden confounding, outcome-related selection, overlap restriction, measurement distortion, and nonlinear response mechanisms, causal pretraining reduced mean precision in estimation of heterogeneous effect (PEHE) by 42.8–53.8% relative to observational pretraining. Second, in a low-dimensional target regime, a directly fitted ridge learner outperformed SciCFM, demonstrating that pretraining is not universally advantageous. Third, in a harder regime with 30 covariates, six sparse active variables, latent response subgroups, and nonlinear threshold, saturation, interaction, and piecewise effects, SciCFM reduced PEHE relative to the strongest direct learner by 29.4% at n=16, 14.6% at n=32, and 3.2% at n=64. These pilot results support a regime-dependent conclusion: causal pretraining is most useful when target samples are extremely small and causal mechanisms are complex but recur across tasks; simple correctly regularized estimators remain preferable when the target mechanism is low dimensional.</p>

Show More

Keywords

causal pretraining target scicfm contain

Related Articles

PORE

About

Connect