Abstract
<p>To assess psychological tests and scales for violations of intergroup comparability – termed Measurement Non-Invariance (MNI) in factor analysis and Differential Item Functioning (DIF) in item response theory – statistical testing has emerged as the dominating approach, but inherently fails to differentiate between major and negligible violations. To address this, upwards of 27 effect size measures for both MNI and DIF have been proposed over the last decades, using vastly different conventions in naming and notation, with no clear consensus on interpretation. This arguably constitutes a severe hindrance for the much needed popularization of a non-dichotomous assessments of intergroup comparability. We address this limitation by introducing a unifying taxonomy that encompasses all previous measures as instances of just four fundamental types: Expected Difference (ED), Expected Absolute Difference (EAD), Total Expected Difference (TED) and Total Expected Absolute Difference (TEAD), which are universally applicable, irrespective of the underlying measurement model. Furthermore, we construct corresponding effect directions and derive formulas to calulate MNI/DIF-induced biases in commonly used mean difference effects, such as Cohen’s d &amp; Cohen’s f. We provide a dual illustration using both FA and IRT models on real-world data.</p>