Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> The ability to assess the importance of detected traffic objects is essential for an autonomous vehicle to make safe driving decisions. The important object identification task aims to reason about the risk and influence of existing entities on the ego vehicle.Existing ego-centric visual methods are constrained by the insufficient spatio-temporal information in the 2D perspective and fail to incorporate textual information, making it difficult to effectively model the interactions among traffic objects and leverage human knowledge for reasoning.To overcome these limitations, we propose a novel important traffic object identification framework by leveraging a Bird's-Eye View (BEV) for global interaction analysis and multimodal contrastive learning to incorporate human knowledge, thereby enhancing dynamic understanding and reasoning ability.Firstly, a <italic>Contrastive Language-Motion Pre-training</italic> (CLMP) model is developed to bridge the gap between dynamic patterns and human insights. This model semantically aligns object motions with textual descriptions through multimodal contrastive learning, enabling the extraction of human knowledge from object motion attributes.Secondly, to achieve the identification of important traffic objects through global interaction modeling under the BEV perspective, a <italic>Interaction-Aware Multimodal Network</italic> (IAM-Network) is proposed. This network integrates motion attributes, scene attributes, and human knowledge to extract and refine spatio-temporal representations of objects, enhancing the reasoning capability for dynamic characteristics of traffic objects.The experimental results demonstrate that the proposed framework outperforms other related works while maintaining real-time performance, and can reasonably identify important objects in different traffic scenarios. </p>

Show More

Keywords

traffic objects human important object

Related Articles

PORE

About

Connect