Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>The principles underpinning the development of human visual perception remain poorly understood. Current computational models of human visual learning still fail at capturing humans' rich semantic recognition abilities. Here, we introduce Embodied Self-Supervised Learning (E-SSL), a class of computational models that harness the rich statistical structure of concurrent visual and motor signals during embodied interaction with the environment for visual representation learning. E-SSL relies on two prominent unsupervised learning principles: temporal consistency and sensorimotor contingencies. When E-SSL is applied to large-scale recordings of human egocentric videos, eye movements and body movements during natural behavior, it produces high-level semantic visual representations in the absence of additional language input or explicit supervision. The learned representations substantially outperform those of previous models of human vision on a range of downstream tasks. We further find that representations learned via E-SSL become more aligned with those of visual cortex of pre-verbal children than those of an established self-supervised model of biological vision. Finally, E-SSL predicts that the onset of locomotion substantially enhances visual perception and suggests that the dichotomy between the ventral and dorsal visual streams may have its origin in the two complementary learning principles of E-SSL. Together, these results shed light on the learning principles and experiences that jointly drive the development of human vision.</p>

Show More

Keywords

visual learning essl human principles

Related Articles

PORE

About

Connect