RESEARCHInvestigateNOW
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
arXiv study shows LLM-generated benchmarks systematically bias results in favor of the model that generated the test set.
Independent coverage
OneBench grouped these reports as coverage of the same underlying development. Open each source to compare the evidence.