Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents
An academic study diagnoses trajectory-level failure modes in long-horizon search agents to evaluate when search effort improves answers.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
An academic study diagnoses trajectory-level failure modes in long-horizon search agents to evaluate when search effort improves answers.
Researchers propose TELLER, a non-intrusive, cross-layer root-cause analysis framework to diagnose LLM inference performance bottlenecks.
Alibaba researchers release Qwen-CUA, a 397B mixture-of-experts model designed for agentic computer use via direct screenshot analysis.
Researchers identify "Solution Hacking," where LLMs pass reasoning benchmarks via invalid shortcuts rather than actual logical reasoning.
Researchers introduced SWE-Touch, a benchmark evaluating how AI coding agents respond to real-time human code modifications in shared workspaces.
Researchers propose frameworks for justifying demographic targets in generative AI audits when prompts leave demographic realization open-ended.
An academic paper analyzes the systemic bias in commercial AI detectors, showing they consistently disadvantage non-native writers.
A research paper introduces Self-Correction Bench, a framework separating model reasoning failures from knowledge deficiencies.
A survey paper on arXiv analyzes how Reward Models (RMs) enhance LLM reasoning via reinforcement learning training signals and inference-time search.
Researchers introduce LatentMAS, a framework enabling LLM agents to collaborate directly via continuous latent space instead of text.
Researchers evaluated Small Language Models (SLMs) against LLMs for multi-turn customer-service QA using synthetic data.
Researchers introduce Logifus, a framework revealing that LLMs fail to solve logical reasoning tasks when the presentation format is obfuscated.
Academic research demonstrates that transformer models implicitly perform adaptive partial pooling, mimicking classical hierarchical regression.
Researchers introduce MAPLE, a method for generating differentially private synthetic data to bypass the compute limits of DP fine-tuning.
Researchers proposed PROClaim, a multi-agent framework using structured, adversarial debate and progressive RAG to verify complex claims.
Researchers propose LangFIR, a method using Sparse Autoencoders to steer LLM output languages using only monolingual data.
Research demonstrates LLMs fail at reliable stochastic sampling, creating a structural vulnerability in LLM-driven agentic systems.
New research demonstrates that high LLM accuracy on reasoning benchmarks does not guarantee faithful execution of step-by-step procedures.
Researchers propose PUPPET, a DPO-based fine-tuning framework that embeds detectable watermarks in LLM outputs without degrading performance.
Researchers find fine-tuning LLMs on a novel language (PyLang) teaches syntax but fails to transfer semantic reasoning and problem-solving.
Researchers evaluate calibration methods for activation oracles that translate LLM internal states into natural language for auditing.
Researchers demonstrate that generative perplexity (gen-PPL), the standard metric for non-autoregressive models, can be easily gamed.
Researchers introduce OmniCSEval, a benchmarking framework evaluating frontier reasoning systems and small models across 1,800 conversations.
Researchers propose a semantic retrieval framework using embedding and indexing pipelines to measure LLM generation novelty against full training corpora.
Researchers propose MENTOR, a self-evolution framework addressing implicit, domain-specific jailbreak risks in finance and education.
Researchers propose SIEVE, a defense framework against indirect prompt injection in LLM agents using selective integrity verification.
Researchers propose RAG strategies to translate natural language queries into database SQL queries and enterprise REST API calls.
Researchers propose an implicit execution tracing method to audit multi-agent AI systems using only the final generated text output.
Researchers propose GraphER, a graph-based enrichment and reranking method to improve multi-source RAG retrieval without agentic latency.
Researchers propose separating evidence extraction from policy execution to solve common failure modes in agentic memory and RAG systems.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion