RESEARCHMonitorNEXT 12 MONTHS
WANDR: A Benchmark for Wide and Deep Research
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
WANDR introduces a 500-task benchmark evaluating autonomous research agents on wide entity discovery and deep, verifiable web investigation.
Open source