RESEARCHInvestigateNEXT 12 MONTHS
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research shows LLM performance rankings flip when generation token budgets vary, with 3–19% of tasks showing accuracy drops at higher caps.
Open source