OneBench
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation | OneBench: AI Insights