A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
New research proposes Direct Corpus Interaction (DCI) for agentic search, guiding fine-grained corpus exploration with relevance estimates.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
New research proposes Direct Corpus Interaction (DCI) for agentic search, guiding fine-grained corpus exploration with relevance estimates.
Research introduces Cognitive Attribution Graphs (CAGE) to improve inline citation generation in long-form LLM outputs, addressing 'attribution ambiguity'.
New research introduces a two-layer evaluation framework to separate language model execution evidence from correctness, exposing how models fail beyond accuracy scores.
INS-ActBench, a new benchmark, evaluates LLMs on professional actuarial tasks requiring auditable, context-grounded, and tool-executable decisions.
Research quantifies the 'tokenizer tax' for Indian languages, showing LLMs incur higher processing costs due to English-centric subword tokenizers.
Research indicates self-improving agents using self-authored tests for verification can show high internal scores while real performance degrades.
Research explores Parallel Autoregressive Decoding (PARD) for block diffusion language models, showing better alignment with left-to-right generation.
CONSISTRE is a new framework designed to improve consistency in document-level relation extraction using large language models by enforcing relational constraints.
New research from arXiv proposes Cross-Attention Calibrated Deduplication to improve RAG system efficiency by identifying and removing redundant data chunks.
Research explores integrating RAG with locally deployed LLMs for regulatory knowledge management, focusing on epistemic reliability without high-end GPUs.
Research identifies a 'blind spot' in AI agent long-term memory systems where retrievers fail to link implicit knowledge to queries.
Research evaluates closed-loop validation-repair for clinical LLMs in healthcare to achieve structured output schema compliance (ICD-10, CPT, HL7 FHIR).
Researchers introduced LEX-EC, a black-box audit framework for zero-shot LLM personality classification, using lexical ablation to distinguish signal from marginal-distribution effects.
New research explores if transformers can dynamically adapt their reasoning strategies (latent algorithm routing) based on input data characteristics.
Researchers propose SINT-Flow, an LLM-based framework with five operators for fully automated, end-to-end schema integration across diverse input tables.
ELMOD, a 2.7B parameter German language model, is introduced for efficient on-device inference using publicly available data and optimized preprocessing.
Researchers introduce D-Score, a spectral statistic computed from hidden activations in a single forward pass, for detecting hallucination in LLMs.
Researchers propose DataOrchestra, a framework for example-specific, adaptive pretraining data processing to improve LLM downstream performance.
New arXiv research proposes a controlled environment and distillation method to improve multi-turn long-horizon planning in foundation model agents.
Semalith v1.4, a 184M-parameter DeBERTa-v3-base classifier, claims state-of-the-art prompt injection detection, general harm, and compliance.
Research finds LLMs frequently provide inconsistent answers to the same factual or mathematical question when rephrased, across 13 models and four benchmarks.
Research proposes Source-Aware Reranking for RAG, incorporating source provenance and credibility as a reliability prior in document retrieval.
MM-ShiftKV introduces a decode-aware KV selection method for multimodal LLMs to reduce memory footprint by optimizing prefill-stage KV caching.
Research introduces HCG-RAG, a method using schema-constrained causal graphs for retrieval-augmented generation to reduce graph size and cost.
Researchers introduce Tokengeist, a multi-turn attribution tracing method for agentic conversations, improving lineage tracking for LLM responses.
Research proposes CRAFT, a method for enterprise coding agents to learn proprietary API schemas and tool use behavior, improving reliability.
STAIF proposes a stage-wise optimization framework for LLMs to better follow complex instructions with multiple constraints, addressing limitations of holistic reward signals.
Research explores how task-adaptation methods like supervised fine-tuning (SFT) impact large language model alignment, including safety and other behaviors.
Researchers propose Imprompt, a new language framework for prompt programming, aiming to simplify the creation of complex LLM tasks and improve usability.
Research proposes Co-Harness, a method to co-optimize LLM agent model parameters and runtime harnesses (prompts, tools, memory) for automated AI research.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion