Abstract
<title>Abstract</title> <p>Expert collapse the convergence of experts to homogeneous representations is the cen tral pathology of Mixture-of-Experts (MoE). This paper establishes a uni ed theoretical frame work (Gradient Pathway Hypothesis), introduces cross-expert gradient alignment (Γ) as a mi croscopic diagnostic, proposes a systematic taxonomy of anti-Gaussian operators, and completes the full pathway from theory → diagnosis → causal veri cation → architectural solution. The Gradient Pathway Hypothesis states that cross-entropy loss propagates an align ment force through routing degrees of freedom. Cross-expert gradient alignment Γ = 2 N(N−1) i 0, synergistic reinforcement), laminar (Γ ≈ 0, independent learning), and negative-aligned (Γ < 0, zero-sum competition). An ef fective Reynolds number Ree = |Γ|/(1 − |Γ|) provides dimensionless stability diagnostics. Anti-Gaussian operators gradient-level (AR loss, ACB), structural (lifecycle death/rebirth), and boundary-condition (expert freezing) systematically counteract this alignment force. Our experimental matrix spans 20+ conditions across two modalities (vision: CIFAR 100, language: TinyShakespeare/WikiText), four architecture families (GPT-2, OPT, Qwen, SmolLM), and a 190× parameter range (36M7B). Key ndings: (1) The routing function con trols the sign of Γ, and routing sparsity controls its magnitude sigmoid activation produces strong turbulent coupling while hard routing approaches laminar; (2) Mosaic Tile hard rout ing with output-space isolation achieves 82.193.0% Full continual learning accuracy with zero forgetting, zero AR loss, and zero deaths; (3) Γ = 0 is rigorously proven under output-space isolation and validated across vision and language modalities; (4) The CL boundary phase dia gram ((α,ntiles ) sweep) reveals near-zero forgetting across the entire phase space with no phase transition boundary; (5) Classical CL methods (L2/EWC/MAS) exhibit 53826× more forget ting than Tile, revealing shared output space as the common root of CL and MoE collapse; (6) 150-epoch CIFAR pre-training reduces forgetting to exactly 0.0 pp; (7) Three-level fully isolated hierarchical MoE achieves the rst trainable-router CL system without catastrophic forgetting; (8) The Γ/Re diagnostic framework enables early collapse detection (Re predictor with 3-σ alerts) and principled architectural optimization (porous walls, boundary-layer suction). The anti-Gaussian operator hierarchy is established: architectural output-space isolation > boundary-condition freezing > training-time gradient perturbation. The natural conclusion 3 the Gaussian Attractor is not MoE’s destiny is validated through a uni ed framework con necting spectral theory, gradient dynamics, ecological parameters, and architectural design.</p>