RESEARCHInvestigateNEXT 12 MONTHS
The Authenticity Gap in Human Evaluation
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research applies economic utility theory to demonstrate flawed assumptions in standard human rating aggregation for language models.
Open source