RESEARCHMonitorWATCHLIST
Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
A new 96-item paired benchmark evaluates LLMs on backtest auditing by measuring flaw recall against clean-control false positives.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.CL — Computation and Language · RESEARCH
- Published by source
- 24 September 2026
- Collected by OneBench
- 25 Sept 2026, 03:01 UK
Stored source excerpt
arXiv:2609.28090v1 Announce Type: new Abstract: Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies.…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.