Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>This paper presents a lightweight and modular framework for language-guided semantic navigation in indoor environments, combining monocular dense simultaneous localization and mapping (MASt3R-SLAM) with modern vision-language models (Gemma 3 and Qwen2.5-VL). Our system uses MASt3R-SLAM as a geometric backbone to generate real-time dense 3D maps from RGB input, while performing open-vocabulary semantic annotation and natural-language reasoning on keyframes. The pipeline consists of three components: monocular dense mapping, keyframe-centric semantic projection, and language-guided goal retrieval with trajectory planning. The focus of this work is on system-level design and integration rather than introducing new SLAM or VLM algorithms. We evaluate the framework in AI2-THOR simulation and on a real TurtleBot3 platform, analyzing trade-offs between mapping quality, latency, and model size. Results show that the proposed system achieves robust dense mapping and practical semantic querying using only monocular RGB, without depth sensors or model fine-tuning, making it suitable for cost-sensitive deployments in assistive and service robotics.</p>

Show More

Keywords

semantic dense mapping monocular framework

Related Articles

PORE

About

Connect