Abstract
<title>Abstract</title> <p>Purpose. Higher education institutions increasingly rely on generative artificial intelligence (GenAI) text detectors to police academic misconduct, and adoption decisions are usually justified by reported accuracy or area under the receiver operating characteristic curve (AUC). This article argues that these metrics are the wrong basis for consequential decisions about individual students, and quantifies the reliability and fairness of detectors when their output is used as evidence. Methods. A transparent signal detection theory model of a GenAI detector is combined with Bayesian base-rate reasoning and Monte Carlo simulation. Detector discriminability and the prevalence of GenAI misuse are swept across ranges anchored to the peer-reviewed empirical literature, including the largest survey of student GenAI misuse to date. A fairness extension models the documented tendency of perplexity-based detectors to score non-native (L2) English writing as more machine-like, and an unequal-variance analysis establishes that the central result is distribution-free. No human-subjects data are used; all datasets are simulation outputs released with the analysis code. Results. Under realistic conditions (AUC near 0.80 and a plausible misuse prevalence of 5 to 15 percent), the positive predictive value of a flag is low: at 10 percent prevalence the probability that a flagged student actually used GenAI is only 22.6 percent at the balanced operating point, so more than three in four flags are wrong. Reaching a 95 percent evidentiary standard while still catching a meaningful share of true cases is unattainable for current detectors: at 10 percent prevalence a detector with AUC 0.95 can catch at most 26.3 percent of true cases without violating that standard. When a threshold is calibrated to a 5 percent false-positive rate on native writers, modelled non-native writers experience an 83.3 percent wrongful-flag rate, a 16.7-fold disparate impact. The barrier is captured by a single distribution-free inequality: safely accusing at 10 percent prevalence under a 95 percent standard requires a positive likelihood ratio (true-positive rate divided by false-positive rate) of at least 171, whereas even a near-perfect detector delivers about 19. A Monte Carlo of a typical deployment yields roughly 2,213 wrongful accusations per 10,000 submissions (95 percent interval 2,133 to 2,293). Conclusion. The problem is structural, not a matter of choosing a better tool. Detection scores should never be the sole or primary evidence of misconduct, thresholds should be reported alongside prevalence-conditioned predictive values rather than accuracy, and institutions should shift resources from detection toward assessment redesign and prevention.</p>