RESEARCHInvestigateNEXT 12 MONTHS
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
New research evaluates LLM-as-a-judge reliability for long-form text outputs, highlighting limitations of short-form evaluation methods.
Open source