Abstract
<title>Abstract</title> <p>The recent advancement of Large Language Models (LLMs) has highlighted the impressive cognitive capabilities of language-only systems. As AI evolves beyond language by integrating multiple modalities, a central theoretical question arises: How does embodiment enhance machine intelligence beyond what language models can achieve? In this review article, we argue that while language inherently encodes a diverse array of embodied experiences, and language-based AI exhibits remarkable proficiency in representation, problem-solving, and communication, its limitations stem from language’s incompleteness in facilitating learning-to-learn, flexible and personalized problem-solving capabilities. Over the long term, advancing AI will require the incorporation of multimodal intelligence to facilitate the creation of novel shared experiences not captured within linguistic representation, as well as enhancing communicative clarity and fostering a sense of co-participation in human-AI interactions. We highlight three crucial avenues for advancing the integration of multimodality in AI systems: (1) Leveraging multimodal representational learning; (2) Adopting hierarchical modulation across modalities inspired by human cognition; and (3) Incorporating autonomous reinforcement mechanisms to endow embodied agents with intrinsic motivations for exploration and continual learning.</p>