RESEARCHMonitorNEXT 12 MONTHS
Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
A study evaluates LLM coding agents on 30 Eurostat tasks, finding code execution success does not guarantee accurate statistical results.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.CL — Computation and Language · RESEARCH
- Published by source
- 22 September 2026
- Collected by OneBench
- 23 Sept 2026, 03:01 UK
Stored source excerpt
arXiv:2609.22222v1 Announce Type: cross Abstract: Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.