Abstract
<title>Abstract</title> <p>Do large language models genuinely self-correct when prompted to review their own outputs, or do they merely perform correction? We investigated this question by subjecting three frontier models ChatGPT/GPT-5.4, Claude 4.6 sonnet, and Gemini 3.1 pro to eight rounds of iterative self-review across three independent repetitions (72 total rounds, temperature = 0). Each model was asked to answer a question, then repeatedly critique and revise its previous response. Despite generating extensive critiques averaging 300 + words per round, the Answer Change Rate across all 63 review rounds was 0% (0/63): no model ever modified its initial answer. We identified 16 distinct phenomena, including three universal patterns Pseudo-Novel Method Generation (fabricating new analytical approaches post-hoc), Critique Drift (shifting focus away from core claims), and Meta-Review Recursion (critiquing the critique process itself), plus 13 model-specific behaviors. These findings reveal what we term "Performative Self-Correction": models generate sophisticated-appearing critical discourse without substantive revision, maintaining initial positions through escalating rhetorical complexity rather than genuine reconsideration. A comprehensive statistical framework, including mixed-effects modeling, mediation analysis, bootstrap confidence intervals, Bayesian Beta-Binomial estimation (Bayes Factor = 40.1, posterior mean correction rate = 1.2%), and network analysis of 15 PSC features, confirmed the robustness of these findings across all models. This challenges assumptions underlying self-correction mechanisms in AI safety frameworks and suggests current prompting-based correction strategies may produce illusions of reliability improvement. Limitations include single-question scope, deterministic sampling, and absence of external ground truth, indicating need for broader empirical investigation into when and whether LLMs can authentically revise their reasoning.</p>