Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Large language models (LLMs) exhibit empirical scaling laws, yet the mechanistic origin of these power laws remains poorly understood. We propose a stochastic multiscale theory of LLM training dynamics, grounded in a formal analogy with nonequilibrium systems such as turbulence and cosmic structure formation, where scaling laws emerge naturally from scale-space transport processes. Central to the theory is a scale-dependent loss $L_r(r,s)$, the held-out loss when usable context is restricted to distance $r$ at training step $s$. Its negative gradient $E_X(r,s) \equiv -\partial L_r/\partial r$ defines a \emph{loss spectrum} in direct analogy with the turbulence energy spectrum. Training dynamics is modeled as a biased random walk with scale-dependent waiting time $\tau_r \propto r^{-\gamma}$ in $r$-space, governed by a Fokker-Planck equation whose exact solution yields an analytically tractable loss spectrum with inertial-range scaling $E_X \propto r^{-(1+\gamma)}$ and an exponential cutoff at an effective scale $r=\xi$. The key exponent $\gamma = \alpha_H/3 = 1/6$ is determined self-consistently from Heaps' law ($\alpha_H \approx 1/2$). Decomposing the total loss into $L_\infty(s)$ (the residual full-context loss, characterizing the \emph{breadth} of the learning frontier) and $\Delta L_r(s)$ (the \emph{depth} of context dependence), we identify a two-phase training dynamics that bifurcates at a critical step $s^*$ into two simultaneous modes: \textbf{Mode A} (downscale compression), with a shrinking effective scale $\xi_A \propto s^{-3/5}$ and growing $\Delta L_r \propto s^{+1/10}$; and \textbf{Mode B} (upscale frontier absorption), with an expanding scale $\xi_B \propto s^{+3/5}$ and decreasing $L_\infty \propto s^{-1/10}$. The theory directly predicts the empirical loss-vs-data scaling exponent $\alpha_D = \gamma/(2-2\gamma) = 1/10$, so that $L \propto D^{-\alpha_D}$. Treating model size $N$ as the effective Langevin time yields $\alpha_N = \lambda\gamma/(2-2\gamma)$, where the asymmetry parameter $\lambda \leq 1$ encodes the fraction of parameters contributing to learning; $\lambda = 3/4$ gives $\alpha_N = 3/40 \approx 0.075$, consistent with the empirical Kaplan value. The full scaling form $L(N,D) = [(N_c/N)^{\lambda} + D_c/D]^{\alpha_D}$ is recovered from the harmonic combination of two independent bottlenecks. The framework yields falsifiable predictions, including a universal loss spectrum $E_X \propto r^{-7/6}$, and provides diagnostics for long-context generalization, data selection, training curriculum, and architecture design.</p>

Show More

Keywords

propto loss scaling training spectrum

Related Articles

PORE

About

Connect