RESEARCHInvestigateNEXT 12 MONTHS
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
A study of 14 reasoning models reveals that test-time compute scaling fails to reduce factual hallucinations in knowledge-intensive tasks.
Open source