RESEARCHInvestigateNOW
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research reveals LLM-as-a-judge evaluation rankings flip across languages, exposing systemic bias in multilingual model benchmarking.
Open source