RESEARCHMonitorNEXT 12 MONTHS
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research shows local open-weight LLMs used as judges achieve internal score consistency without agreeing with human evaluation benchmarks.
OneBench interpretation
Institutional assessment
So what
Automated LLM evaluation pipelines using smaller open-weight models can pass internal consistency checks while producing fundamentally flawed assessment scores.
Do what
Review model validation protocols with the team responsible for automated testing before deploying LLM-as-a-judge pipelines.