RESEARCHInvestigateNEXT 12 MONTHS
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Academic research demonstrates that LLM-as-a-judge evaluation frameworks fail to reliably assess accuracy in high-stakes domains without human grounding.
Open source