Abstract
<p>Cronbach's alpha is the most commonly reported metric of reliability in psychological measurement, and statistical software encourages a common post hoc revision: dropping the item whose removal most improves in-sample alpha. A Monte Carlo simulation examined what this practice does to true reliability, varying items per scale (5 to 40) and sample size (20 to 1,000), with item pools generated from a loading distribution fitted to 3,783 factor loadings from real psychological scales. Two strategies, always dropping the nominated item or dropping only when in-sample alpha improves, were evaluated against the population alpha and true reliability of the retained items. The bias in the reported alpha itself was modest, overstating true reliability only in short scales at small samples and understating it elsewhere. The decision was another matter: at the smallest scale and sample size, dropping reduced true reliability in 46.7% of the replications in which the conditional rule acted, despite requiring an apparent improvement, and almost the entire apparent gain was selection optimism. Stricter decision rules helped only in large samples: with N of 200 or more, acting only on apparent gains above .02 kept the probability of harming the scale at or below 7%, whereas at N = 20 no threshold examined made the decision safe. Because this risk depends on scale length and sample size, given loadings typical of trait scales, it is knowable before data are collected, whereas nothing inside the alpha-if-item-removed table distinguishes a genuinely weak item from a sampling artefact.</p>