Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>The Generalized Advantage Estimator fixes a single bias–variance tradeoff (λ) for every state and training phase. We introduce UGTC, a plug-in module that replaces the advantage estimator with a state-dependent blend of a fast critic (λ=0.80) and a slow critic ensemble (λ=0.99, M=3), gated by ensemble disagreement. UGTC composes with PPO, TD3, SAC, and DreamerV3 via backbone-specific insertion points. Across six benchmarks (64 tasks, 10 seeds, bootstrap 95% CIs), UGTC-PPO converges 2.7× faster on Hopper (1.9× wall-clock); UGTC-DreamerV3 peaks at 466.7 vs. 191.8 on MetaWorld ML45 (2.4×); UGTC-PPO improves Procgen Hard by +19.9%. We report two negative results—Crafter (−23.3%) and Ant peak return (−14% vs. TD3)—and provide deep analysis of both. A systematic ablation including meta-gradient λ, REDQ, a learned gate, and a per-state λ network confirms that calibrated epistemic uncertainty is the critical ingredient. A held-out sensitivity analysis on Procgen and ML45 (not the tuning benchmark) validates hyperparameter robustness.</p>

Show More

Keywords

advantage estimator ugtc critic ensemble

Related Articles

PORE

About

Connect