Abstract
<jats:p>We develop a parametric simulation framework that models the training dynamics of two self-training architectures: a classical pseudo-labeling baseline and a state-of-the-art (SOTA) pipeline combining RLAIF (Reinforcement Learning from AI Feedback), Self-Instruct generation, and multi-agent consensus filtering. Each observable of interest — validation accuracy, loss decay, per-domain F1-Score, and curation latency — is expressedas a closed-form function of the training epoch or batch size, and the free parametersare calibrated to reflect the qualitative behaviour reported for each architecture. Bynumerically integrating these models we quantify the convergence ceiling, optimisationstability, domain-wise accuracy, and throughput scaling of the two pipelines. The simulationshows that consensus-filtered curation raises the accuracy ceiling by roughly seven points,suppresses the residual loss by an order of magnitude, and keeps curation latency sub-linearin batch size. This work is a modeling and simulation study; all curves are generated fromthe stated equations rather than measured on a physical training run.</jats:p>