RESEARCHMonitorWATCHLIST
Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
A study across six LLMs shows model accuracy predicts bias in absolute scoring tasks, impacting LLM-as-a-judge reliability.