Abstract
<title>Abstract</title> <p>Multinomial logit estimation can be unstable when predictors are nearly collinear and can be distorted by response miscoding and high-leverage observations. We develop a cross-fitted contamination-aware generalized empirical-Bayes Liu estimator (CF-CABLS-MNL) that combines weighted density-power-divergence estimation, a diagonal-dominant class-misclassification model, sandwich/Godambe geometry, and direction-specific spectral shrinkage. Out-of-fold robust fits provide both a prior centre and clean-class probabilities for a MAP–EM estimate of the misclassification matrix, thereby avoiding the direct reuse of one full-sample coefficient estimate as both prior centre and data update. The final robust coefficient estimator is rotated into the eigenspace of the estimated Godambe precision, where eigenvalue-adaptive beta priors determine direction-specific Liu factors and are integrated by deterministic quadrature. The empirical study contains 16 Monte Carlo scenarios, eight estimators, 100 paired replications per scenario, a 30-replication sensitivity study, and five repeated semi-synthetic train–test experiments on the UCI Dry Bean, Vehicle Silhouettes, and Glass Identification data. Ridge regression minimized coefficient error and predictive log loss in all 16 simulation scenarios, emphasizing the strength of direct regularization when the data-generating coefficient vector is dense and the ridge constant is favorable. Relative to maximum likelihood, however, CF-CABLS reduced average SMSE by about 19%, and the version without the misclassification matrix reduced it by about 30%. Full CF-CABLS achieved the best balanced accuracy in six scenarios, especially under joint response and leverage contamination. In the real-data experiments no estimator dominated uniformly: CF-CABLS achieved the lowest log loss for Dry Bean under combined contamination, whereas MLE, DPD, scalar Liu, and ridge were preferable in several clean or small-sample settings. All reported fits completed without recorded optimization failures. Spectral diagnostics showed a strong positive association between robust-information eigenvalues and posterior Liu factors, confirming that poorly identified directions received stronger shrinkage. The evidence supports CF-CABLS as a transparent robustness–shrinkage framework and ablation device, while also showing that explicit misclassification adjustment should be used selectively rather than assumed to improve every dataset.</p>