Agentic Routing: The Harness-Native Data Flywheel
Research explores 'agentic routing,' where an execution harness manages specialized LLM agents for different tasks, optimizing for model strength.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research explores 'agentic routing,' where an execution harness manages specialized LLM agents for different tasks, optimizing for model strength.
New research proposes Epistemic Asymmetry Schelling Task (EAST) to evaluate LLM Theory of Mind, moving beyond easily 'gamed' traditional tests.
RefineEvo is a new evolutionary framework for Automatic Heuristic Design (AHD) that addresses LLM-based method limitations in combinatorial optimization.
Research introduces AI Textbook Auditor, a multi-agent LLM system for automated quality assurance of educational materials, assessing factual accuracy, technical correctness, and linguistic quality.
Research paper explores how generative AI changes the validity of high-stakes English language proficiency tests, suggesting 'construct drift'.
Research explores using inference-time feed-forward network interventions to improve structured outputs and tool-use by LLMs without retraining.
Researchers introduced a new dataset and benchmark for query-focused summarization, specifically for event-oriented thematic document corpuses.
Research proposes Unified Gradient Projection (UGP) to prevent catastrophic forgetting in multilingual ASR models fine-tuned on low-resource languages.
TIGER proposes a new multimodal speculative decoding method using text-conditioned visual gated routing for improved VLM acceleration.
Research finds freely accessible LLMs fabricate legal citations for GDPR and Saudi data protection law, using a new bilingual benchmark.
ResearchQA introduces a new benchmark with 6,211 question-answer pairs across scientific papers for citation-grounded LLM evaluation.
LLMs struggle with pragmatic cooperation when collaborators have asymmetric information, according to new research investigating their agentic capabilities.
Research used GPT-4.1 to decompose customer satisfaction in 9,000 support conversations, validating LLM annotations against self-reported scores.
Research explores whether large language models can continually acquire and retain new facts by writing them directly into model weights without catastrophic forgetting.
Research proposes a framework using LLMs for token-efficient and semantic-preserving opinion summarization from large, redundant text corpora.
Research identifies "Thinking Collapse" in On-Policy Self-Distillation (OPSD) for LLMs, causing performance degradation in complex reasoning tasks.
New research introduces QIMG-7, a benchmark for evaluating multimodal RAG systems against polluted retrieval from unreliable sources, addressing false text and misleading images.
Researchers introduced FATE, an 8B-parameter model for evaluating AI tutors, addressing the lack of reliable methods for LLM pedagogical quality.
MafiaScope is an open testbed using the game Mafia to measure LLM agents' Theory of Mind through structured probe questions after public utterances.
New Eval-Pair Matrix protocol improves meta-evaluation of LLM-as-a-judge for RAG by identifying self-leniency and answer-causal contradictions.
Research finds that including more demographic attributes in LLM prompts can decrease agreement between LLM predictions and human annotations.
UNIBROWSE introduces a data-to-agent framework for multimodal browsing agents to handle complex web content and information flows.
Research introduces Kahneman4Review benchmark to evaluate LLM judges on epistemic reliability in peer review, differentiating LLM and human quality assessments.
Research finds Chain-of-Thought (CoT) prompting can hinder LLM performance in legal drafting, specifically patent claim generation.
Research introduces "Diversion Decoding," a method to improve hallucination detection in LLMs by identifying factually incorrect generations.
New research introduces EYT-Bench, a human-centered benchmark for multi-turn dialogue LLM evaluation, using a decoupled user simulator and target model.
Research proposes a multi-agent framework called 'expert-editor stepwise questioning' to improve long-document summarization by LLMs.
Researchers introduced 'Structured Thoughts,' a framework for LLMs to organize reasoning into alternating exploratory and distilled conclusion blocks to improve efficiency.
CAFE is an open-source framework applying design of experiments to evaluate compound AI systems by systematically testing swappable components.
Research presents a hybrid model using AfroXLMR-Social and DeBERTa for detecting polarized discourse in English and Hausa across resource settings.
CLIR-Bench is a new benchmark for multimodal question answering over sparse, irregularly sampled, and asynchronous clinical time series data.
Bilibili's Index-1.9B is a new series of 1.9B parameter open small language models, pre-trained on 2.8T predominantly Chinese and English tokens.
RouteRec framework proposes strict evaluation of recommender-agent selection and aggregation across traditional and LLM-based rerankers.
Research introduces a benchmark framework for evaluating the faithfulness of LLM-generated clinical trial summaries, addressing hallucination risks for multi-stakeholder audiences.
Research finds post-training quantization can silently alter LLM reasoning and introduce new failure modes, even when accuracy metrics remain stable.
FindMyText is an open-source Python package for robust, scalable detection of text containment and near-verbatim copies within large corpora.
Research indicates LLMs exhibit weaker reasoning capabilities and higher inference costs when processing non-English languages like Japanese.
Research paper introduces PaperRouter-Agent, an LLM agent for personalized hierarchical document routing into user-defined folksonomies.
Research finds search API outputs (snippets vs. full pages) significantly impact agent decision-making, even with equal underlying accuracy.
Research explores LLMs' ability to dynamically select preference aggregation strategies in group recommenders based on fairness perceptions.
Research explores how language models can handle different 'mental spaces' like beliefs or hypotheticals, suggesting a shared router mechanism.
Research explores using language similarity for cross-lingual transfer in automatic speech recognition (ASR) for low-resource languages like Warlpiri.
Research investigates how information locality in language affects LLM's ability to reconstruct natural language from syntactically disrupted input.
Research introduces 'Weaver,' a new speculative decoding method improving token generation efficiency and acceptance rates for autoregressive models.
Research from arXiv proposes STEP, a model for career-path recommendation using temporal and educational trajectory modeling from resumes.
A research paper proposes a human-centered AI framework for AI-assisted lexicography, addressing concerns about the role of lexicographers and linguistic diversity.
Research introduces Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to measure LLM performance variance due to prompt formatting.
Researchers propose ISE, a three-stage synthetic data generation paradigm (Intent -> Simulate -> Execute) for training multi-turn OS agents.
Research proposes an evaluation-unsupervised method for selecting small subsets of prompts from LLM benchmarks to approximate full benchmark scores.
Research explores enforcing norms on AI agents in multi-agent systems to prevent self-serving behaviors that harm collective goals, similar to human societies.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion