Abstract
<title>Abstract</title> <p>We audit gender bias in legal large language models (LLMs) by choosing a setting where ground truth is unambiguous: wrongful-dismissal compensation under Article 87 of China's Labour Contract Law, whose value is uniquely determined by a statutory formula. Across 400 deterministic inferences on four mainstream LLMs, we document a form of bias that prior work has largely overlooked: counterfactual instability under deterministic decoding. All four models violate counterfactual fairness (mean |Δ| = 0.082–0.266), yet within our sample no statistically significant directional bias is detected (Wilcoxon p > 0.28); the distribution of |Δ| is heavy-tailed, with median values of zero for two models alongside extreme individual deviations. This pattern matters because case-specific instability, unlike directional bias, cannot be corrected by any aggregate statistical adjustment and is invisible to group-level audit approaches. Targeted at this failure mode, we validate two complementary remedies: Counterfactual Symmetric Calibration (CSC), a domain-knowledge-driven post-processing safety net that raises legal compliance by 14–72 percentage points at near-zero cost; and Fairness-Constrained Fine-Tuning (FCFT), a diagnostic fine-tuning framework whose ablation identifies missing formula knowledge, rather than encoded gender preference, as the primary driver of biased outputs. Together, these results provide a reproducible auditing-to-remediation pipeline and a principled argument for decoupling legal calculation from language generation in AI-assisted adjudication.</p>