RESEARCHInvestigateNEXT 12 MONTHS
LLM Evaluators are Biased across Languages
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
LLM evaluators (reward models and LLM-as-a-Judge) exhibit significant language bias, failing to assign consistent scores across 23 languages for semantically identical instruction-response pairs.
Open sourceOneBench interpretation
Institutional assessment
So what
Multilingual model deployments require validation frameworks that account for known language biases in LLM evaluators, challenging current assumptions of language-neutral scoring.
Do what
Your model validation framework for global LLM deployments needs to explicitly address cross-language evaluation bias and cannot rely solely on pairwise accuracy metrics.