RESEARCHInvestigateNEXT 12 MONTHS
Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Autorubric introduces a framework to mitigate position bias, conflation, and calibration errors in LLM judges for non-verifiable tasks.
Open source