RESEARCHInvestigateNEXT 12 MONTHS
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
A forensic audit of a radiology VLM benchmark found inconsistencies across datasets, DICOM rendering, prompts, APIs, and statistical code artifacts.
Open sourceOneBench interpretation
Institutional assessment
So what
Benchmarking inconsistencies in even preserved VLM pilots compromise model reliability and validation efforts.
Do what
Add VLM benchmark auditability to your model validation framework requirements.