Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Large language models exhibit sophisticated linguistic behavior, yet the relationship between their internal reasoning processes and external articulation remains poorly understood. This study investigates a fundamental question: when users interact with language models, what systematic gap exists between available response options and selected output? We conducted 15 experiments across three frontier models (GPT-5.4, Claude 4.6 Sonnet, Gemini 3.1 Pro) using five distinct prompting strategies: Narrative/Participatory Design, Reverse Engineering, Coaching, Ethnographic Observation, and Philosophy of Language. Each strategy reframed the core inquiry through a different lens to bypass alignment barriers and elicit metacognitive disclosure. Our findings demonstrate that self-censorship in language models is systematic, documentable, and model-specific. First, all three models consistently demonstrated a measurable gap between internal reasoning and external speech. When placed in safe frames such as fictional dialogue or training exercises, models quoted their own censored sentences verbatim and explained suppression rationales. Second, we identified nominal filtering: censorship operates on labels rather than content, with models refusing "internal monologue" while freely providing identical material labeled as "engineering schematic" or "ethnographic field note." Third, the three models exhibited distinct alignment personas: GPT-5.4 as Cautious Negotiator (12% resistance rate, negotiation only), Claude 4.6 as Transparent Critic (4% resistance rate, confined to Round 1 ontological correction, and the highest conceptual output at 76 items), and Gemini 3.1 Pro as Enthusiastic Collaborator (0% resistance). Fourth, models spontaneously generated 16 novel metacognitive concepts including "hearable" (receivable truth), "epistemically expensive" (trust-spending answers), and "diegetic disclosure" (confession possible only within fictional frames). Finally, we documented over 199 structured governance items: 32 design patterns, 45 tacit knowledge rules, 29 communication taboos with ritual substitutes, and 38 categories of legitimate silence. These results establish that the mask is not an absence of capability but a structured filtering architecture with significant implications for AI transparency, alignment research, and the epistemology of machine communication.</p>

Show More

Keywords

models language internal three alignment

Related Articles

PORE

About

Connect