Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Vision–language segmentation models are increasingly deployed as if they understood referring expressions the way humans do. Yet their evaluation is dominated by photographic benchmarks where lexical shortcuts can mask whether a model has parsed the prompt. We isolate, measure, and characterise a failure mode we call contextual amnesia: the systematic collapse of state-of-the-art referring segmentation systems on prompts requiring relational reasoning, multi-target selection, instruction negation, or empty-target rejection, despite competent attribute grounding. We design a controlled synthetic benchmark of 10000 images and 52791 rendered objects, evaluate six open-source pipelines across five linguistic regimes on 105 metrics, and report a 4–6× gap between attribute and relational prompts and a complete 0.00 IoU collapse on negation and empty-target prompts for all detector-cascade pipelines. Three ≈9M-parameter randomly-initialised architectures trained from scratch surpass the strongest billion-parameter cascade by +24 to +25 IoU points. Five targeted adaptation techniques, and a late-fusion ensemble of two specialists, reach 0.687 IoU—2.00× the strongest zero-shot pipeline.</p>

Show More

Keywords

prompts segmentation referring collapse relational

Related Articles

PORE

About

Connect