RESEARCHInvestigateNEXT 12 MONTHS
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers introduced SABRE, an automated pipeline that generates stress tests for Vision-Language Models to identify model weaknesses.
Open source