RESEARCHMonitorNEXT 12 MONTHS
LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers introduced LoopsBench, a long-horizon benchmark evaluating coding agents on sequential tasks structured as dependency DAGs.
Open source