RESEARCHInvestigateNOW
The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research shows vision-language models ignore visual inputs on up to 97% of benchmark samples, relying instead on textual language priors.