OneBench
Clinician use of language models diverges from how the models are evaluated | OneBench: AI Insights