OneBench
The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World | OneBench: AI Insights