RESEARCHInvestigateNEXT 12 MONTHS
The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research demonstrates that offline log replay fails to accurately evaluate model switching inside multi-step agentic workflows.
Open source