RESEARCHInvestigateNEXT 12 MONTHS
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research demonstrates fundamental scaling limits and systematic bias in LLM-as-a-judge methods compared to traditional data annotation.
Open source