Abstract
<p>Deep learning models for face recognition are widely adopted in cognitive neuroscience, yet it is not well understood whether their excellent performance means they “see” faces like humans. Here, we collected a substantial dataset of human face-similarity judgments to assess how well prominent face recognition models align with human perception. We found that models with high—but not the highest—recognition performance are often best aligned with human similarity judgments. We further tested to what extent a linear transformation could better align model representations with human perceptual similarity. Models with superior recognition performance benefited the least from this transformation, suggesting a deeper mismatch with human perception. Notably, enforcing this alignment reduced recognition performance for higher-performing models but gave low-performing models a slight recognition boost. Together, these results indicate that recognition ability and human-like perceptual similarity, although related, involve a trade-off beyond a certain level. This trade-off was observed not only across models but also across processing stages within models, further supporting the generality of this relationship. This work may inform the development of more human-like models and the use of deep learning models for studying human perception.</p>