RESEARCHInvestigateNEXT 12 MONTHS
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research reveals LLM-as-a-judge systems predict scores using rubric text alone without reading responses, uncovering rubric leakage artifacts.
Independent coverage
OneBench grouped these reports as coverage of the same underlying development. Open each source to compare the evidence.