Abstract
<jats:p>Mesoscale eddies modulate the regional and seasonal variability of ocean biogeochemistry, but their chaotic dynamics make individual eddies unpredictable beyond a few weeks. Large ensembles of high-resolution models can reproduce these mesoscale statistics, but are computationally prohibitive. We present SamudraBGC, a machine-learning emulator trained on an idealized double-gyre North Atlantic configuration of a 9 km ocean biogeochemical model. SamudraBGC reproduces the contrasts between the oligotrophic subtropical gyre, the productive subpolar gyre, and the jet, and the seasonal to interannual variability. Four choices mattered most: a Helmholtz velocity decomposition to produce coherent eddies, log-transformation of biogeochemical tracers to capture their extremes, a gradient loss to preserve fronts, and vertical compression to improve fidelity at depth. A 50-member ensemble broadly reproduces the spread and growth rates of the dynamical ensemble -- using $\sim$0.1 GPU-hours per model year against $\sim$3,600 CPU-hours for the dynamical model -- paving the way towards ensemble-based data-assimilation of ocean biogeochemistry.</jats:p>