RESEARCHInvestigateNEXT 12 MONTHS
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Diagnostic evaluation of AI agents on 100 frontier research tasks identifies systematic operational failure modes in autonomous workflows.
Open source