Back to Search View Original Cite This Article

Abstract

<p>Value-aligned multimodal large language models (MLLMs) appear explicitly unbiased, yet they still encode implicit biases. Quantifying implicit biases is essential for pursuing fairness in MLLMs. However, measuring them remains an open challenge. Here, we propose reverse correlation, a widely used data-driven technique in psychology for capturing mental representations and stereotypes, as a method for uncovering implicit biases in MLLMs. Applying this technique, we investigated whether state-of-the-art MLLMs, GPT-5.4 and Claude-Opus-4.8, reflect stereotypical gender biases in visual templates of leaders and followers reconstructed from 75,600 choices. We find that human participants perceived leader templates as more masculine and follower templates as more feminine. We further demonstrate that Claude-Opus-4.8-generated templates showed stronger stereotypical gender associations than human representations, indicating amplification of gender bias. Our findings advance implicit bias evaluation in MLLMs by establishing reverse correlation as a tool for uncovering visual stereotypes and highlighting the need to mitigate implicit biases despite explicit alignment efforts.</p>

Show More

Keywords

mllms implicit biases templates gender

Related Articles

PORE

About

Connect