RESEARCHInvestigateNEXT 12 MONTHS
Noise Floor Audit for Agent Benchmarks
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
An arXiv audit of tool-calling agent benchmarks reveals semantic prompt perturbations create significant score variance across provider endpoints.
Open source