RESEARCHMonitorNEXT 12 MONTHS
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research identifies Rubric-Induced Preference Drift, where minor edits to LLM-judge rubrics systematically alter model evaluations.
OneBench interpretation
Institutional assessment
So what
Flawed LLM evaluation rubrics can stealthily bias model selection and validation controls while successfully passing standard benchmark checks.
Do what
Review validation controls for automated LLM judge rubrics with the team responsible for model risk management.