Abstract
<p>Fairness evaluation in automated essay scoring (AES) requires more than determining whether group differences exist. It also requires evidence about how large those differences are, which group they favor, and whether they warrant further attention. This study introduces two model-level indices, Conditional Mean Disparity (CMD) and Conditional Odds Disparity (COD), for evaluating group differences in predicted writing scores after accounting for human-rated proficiency. CMD quantifies disparity on the predicted score scale, whereas COD quantifies disparity using the relative odds of receiving higher predicted scores. The indices were demonstrated using writing samples from the PERSUADE 2.0 corpus and four AES systems representing traditional statistical modeling, deep learning, and prompt-based large language model scoring. The findings showed that CMD and COD provided interpretable estimates of disparity magnitude across systems and captured information that was not reflected in scoring performance alone. By emphasizing the size and direction of observed group differences, the proposed indices can support fairness evaluation, model comparison, and decisions about whether further examination is needed.</p>