RESEARCHMonitorNEXT 12 MONTHS
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
DSAgentBench proposes a benchmark evaluating autonomous AI agents on end-to-end data science workflows in real computer environments.
Open source