Abstract
<title>Abstract</title> <p>Multimodal conversational emotion recognition aims to identify utterance-level emotions by integrating textual, acoustic, and visual affective signals. Existing methods have improved contextual modeling and multimodal fusion, but they mainly focus on global representation aggregation and pay insufficient attention to ambiguous local decisions between closely competing emotions. To address this issue, we propose Evidence-Aware Multimodal Fusion (EAMF), a decision-oriented framework for multimodal conversational emotion recognition. EAMF obtains a base multimodal prediction and constructs complementary evidence from dialogue-state retrieval, cross-modal conflict comparison, prototype-based boundary modeling, and trigger-aware temporal cues. Rather than revising all emotion classes uniformly, EAMF focuses on the two most competitive emotion classes. State and conflict evidence are used for pair-wise margin calibration, while boundary and trigger evidence provide auxiliary boundary supervision and contextual state modulation. This design preserves non-competitive classes and supports evidence-aware local decision modeling for difficult samples. Experiments on IEMOCAP and MELD show that EAMF achieves improved overall performance on IEMOCAP and competitive results on MELD, especially for several challenging minority categories. Further analyses support the effectiveness and complementarity of the proposed evidence-aware modeling strategy.</p>