RESEARCHInvestigateNEXT 12 MONTHS
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers evaluate verbalized confidence in small language models (0.5B-14B) to determine if they can safely defer to human reviewers.
Open source