RESEARCHInvestigateNOW
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
ArXiv study reveals agent benchmarks measure task specialization rather than capability, with agent choice causing under 3% of variance.
Open source