OneBench
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure | OneBench: AI Insights