RESEARCHInvestigateNEXT 12 MONTHS
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research introduces AgentRelBench, demonstrating that single-run evaluations miss catastrophic non-deterministic agent failure modes.
Open source