RESEARCHInvestigateNEXT 12 MONTHS
Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Academic study evaluating 23 LLMs reveals that standard commonsense benchmarks fail to reliably predict performance on downstream tasks.
Open source