OneBench
Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors | OneBench: AI Insights