Abstract
<jats:p>The most accurate neural decoder on held-out trials is not necessarily the most useful for brain-computer interfaces or neural population analysis. In practical use, neural decoders may also need to remain robust to noisy neural inputs, satisfy calibration or deployment constraints, and produce comparable representations across recordings. We introduce BEND-BCI, an open-source benchmark of 23 neural decoding methods on motor, visual, speech and spatial decoding tasks across 16 real or synthetic neural recordings. BEND-BCI compares decoders across held-out prediction, robustness to input perturbation, computational cost and cross-recording latent consistency. These additional axes frequently changed decoder rankings: held-out accuracy did not reliably identify the most robust, efficient or cross-recording-consistent models. Simpler baselines were also competitive with, and in some cases outperformed, more heavily parameterized deep neural networks. Diagnostic analyses based on explainable machine learning further showed that decoder performance was associated with the use of expected neural features and could be improved by selecting high-quality training trials. BEND-BCI reframes neural-decoder selection from an accuracy leaderboard into a constrained decision over task, resource, representation and diagnostic goals.</jats:p>