RESEARCHInvestigateNEXT 12 MONTHS
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research paper shows outcome-only LLM agent evaluation misses execution path failures, policy violations, and inefficiencies in tool usage.
Independent coverage
OneBench grouped these reports as coverage of the same underlying development. Open each source to compare the evidence.