Back to Search View Original Cite This Article

Abstract

<jats:p>We develop a parametric simulation framework that models the training dynamics of two self-training architectures: a classical pseudo-labeling baseline and a state-of-the-art (SOTA) pipeline combining RLAIF (Reinforcement Learning from AI Feedback), Self-Instruct generation, and multi-agent consensus filtering. Each observable of interest — validation accuracy, loss decay, per-domain F1-Score, and curation latency — is expressedas a closed-form function of the training epoch or batch size, and the free parametersare calibrated to reflect the qualitative behaviour reported for each architecture. Bynumerically integrating these models we quantify the convergence ceiling, optimisationstability, domain-wise accuracy, and throughput scaling of the two pipelines. The simulationshows that consensus-filtered curation raises the accuracy ceiling by roughly seven points,suppresses the residual loss by an order of magnitude, and keeps curation latency sub-linearin batch size. This work is a modeling and simulation study; all curves are generated fromthe stated equations rather than measured on a physical training run.</jats:p>

Show More

Keywords

training accuracy curation simulation models

Related Articles

PORE

About

Connect