Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
Factual evidence
What the source reports
Research reveals fixed LLM judges make version-dependent errors when evaluating updates to coding and customer-service agents.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.LG — Machine Learning · RESEARCH
- Published by source
- 30 September 2026
- Collected by OneBench
- 1 Oct 2026, 03:02 UK
Stored source excerpt
arXiv:2609.34198v2 Announce Type: replace Abstract: Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.
OneBench interpretation
Institutional assessment
So what
Automated LLM-as-a-judge evaluation frameworks can produce systemic errors when testing iterative enterprise AI agent upgrades.
Do what
Review model validation protocols for automated LLM evaluation pipelines to ensure version-dependent bias is accounted for.