RESEARCHInvestigateNEXT 12 MONTHS
Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers demonstrate that multimodal model confidence readouts can be manipulated via image attacks while preserving the exact text answer.
Open source