RESEARCHInvestigateNEXT 12 MONTHS
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
arXiv research reveals LLM judges show bias when evaluating functionally correct code that has superficial formatting or style variations.
Open source