Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Vision-language models (VLMs) have demonstrated strong zero-shot learning (ZSL) capability in open-category recognition through joint image-text modeling. However, existing approaches for improving zero-shot performance often rely on large-scale pretraining, prompt learning, or lightweight adaptation, which involve high training costs, strong task dependence, or insufficient optimization of the shared representation space. In particular, during post-pretraining, the modality gap, distributional degradation, and representation collapse may emerge simultaneously, limiting further improvements in zero-shot classification performance. To address these issues, this paper proposes GeoRel-CLIP, a geometry-regularized relational distillation method for CLIP post-pretraining. Instead of solely reducing the distance between image and text features, GeoRel-CLIP jointly optimizes the shared embedding space from two perspectives: global geometric distribution and local relational structure . Specifically, we introduce a Geometric Distribution Regularization (GDR) module to regulate feature spreading and directional occupancy balance on the unit hypersphere, thereby suppressing excessive feature aggregation and representation collapse. Meanwhile, we further propose a Scale-Invariant Relational Distillation (SIRD) module, which standardizes and aligns the cross-modal relational matrices of the teacher and student models, reducing the interference caused by scale drift and enhancing the consistency of structural knowledge transfer. Experimental results show that GeoRel-CLIP outperforms existing SOTA methods on multiple standard zero-shot datasets and achieves further improvement in the average Top-1 accuracy across 11 datasets.</p>

Show More

Keywords

zeroshot relational representation further georelclip

Related Articles

PORE

About

Connect