RESEARCHInvestigateNEXT 12 MONTHS
Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Study reveals expert and crowd annotator pools disagree on 23.6% of benchmark preference data, exposing flaws in standard model evaluations.
Open source