Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> <bold>Background</bold> : Local objectives expose parallelism across network depth by removing cross-stage backward dependencies, but they alter end-to-end credit assignment and complicate ownership of a shared output readout. <bold>Objectives</bold> : We evaluate the quality, throughput, and memory trade-off of bounded-horizon local training for 24-layer byte-level Transformers on a many-core CPU, and investigate how upstream local losses should participate in shared-readout updates. <bold>Methods</bold> : We introduce readout-gradient consensus (RGC), in which upstream stages differentiate local losses through same-microbatch readout snapshots and the final stage averages their contributions. Six matched methods are evaluated under a prospectively specified protocol using paired seeds, held-out testing, repeated systems measurements, two corpora, and three model widths. <bold>Results</bold> : On a 31-core allocation, asynchronous RGC reaches 1.382 times the throughput of matched global backpropagation, with a paired 95% interval from 1.364 to 1.400, while peak process-tree proportional set size increases from 1.94 to 4.31 GiB. The comparison uses 29 stage-worker threads versus 24 intra-operation threads for backpropagation. The selected asynchronous broadcast method has a held-out mean quality gap of 0.841%, but its one-sided 95% upper bound is 2.095%; therefore, the prespecified 1% non-inferiority criterion is not met. Additional enwik8 widths satisfy their joint criteria, whereas TinyStories does not, and readout equivalence is not established. <bold>Conclusions</bold> : The evidence supports a bounded CPU quality-throughput-memory frontier for this implementation, not quality-preserving training or a general readout-recovery mechanism. </p>

Show More

Keywords

local readout objectives quality throughput

Related Articles

PORE

About

Connect