OneBench
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification | OneBench: AI Insights