RESEARCHInvestigateNOW
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research reveals LLM confidence estimation techniques are highly sparse, with models like Qwen3-32B clustering most outputs at exactly 95%.
Open source