Abstract
<jats:p>Composed image retrieval (CIR) is a complex image retrieval task that requires using both a reference image and a corresponding caption in the query (Zhang et al., 2025). The reference image provides visual context, while the caption specifies the desired modification. Together, they form a composed query that captures user intent more precisely than image or text alone. This makes CIR particularly valuable in fashion retail, where users frequently seek visually similar items with specific attribute changes. While monolingual CIR systems built on English particularly CLIP-based models have shown strong benchmark performance, multilingual and cross-lingual composed image retrieval remains substantially underexplored. Existing fashion CIR systems operate exclusively in English, leaving over 200 million Urdu speakers and over one billion Chinese speakers underserved. Explainability has also been largely ignored, with no existing fashion CIR system offering interpretable retrieval decisions. This research addresses both limitations through a multilingual explainable CIR framework that supports English, Urdu, and Chinese queries within a single unified pipeline, built on the FashionIQ benchmark. The framework is evaluated on FashionIQ comprising three clothing categories with over 32,000 training triplets using Recall@K as the primary evaluation metric across all three query languages. It integrates a cross-modal projection head for multilingual embedding alignment, a learnable weighted fusion module, LLaVA-1.6 for query enrichment, and Grad-CAM for visual explainability. This is a novel CIR framework that jointly addresses multilinguality and explainability in a unified pipeline, with the proposed model evaluated at each stage of the pipeline from multilingual query specification through to visually explained retrieval results.</jats:p>