Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks
Research identifies a benchmark leakage issue where 'out-of-distribution' classes were inadvertently included in training, skewing OOD detection metrics.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research identifies a benchmark leakage issue where 'out-of-distribution' classes were inadvertently included in training, skewing OOD detection metrics.
New research, HyGRL, proposes a hybrid graph reasoning framework for multi-entity questions, addressing limitations in conventional RAG and LLM-constructed Graph-RAG.
Researchers propose JailMeter, a new evidence-based framework to consistently evaluate jailbreak attack effectiveness on large language models.
New research proposes "knowledge-centric self-improvement" for AI agents, shifting focus from optimizing agent design to improving knowledge bases for better transferability and maintenance.
Research introduces 'Twin Agent', a method for secure LLM agents using context residual compression to mitigate prompt injection risks.
New research proposes methods for LLM preference alignment that reward reasoning trajectories, not just final outcomes, addressing coarse credit assignment.
Research introduces specialist language models for high-precision, scalable missing value prediction in tabular data, addressing overconfidence and hallucination.
DocOps introduces a deterministically verifiable evaluation framework for autonomous agents focused on complex document operations, deconstructing tasks into atomic dimensions.
Researchers propose Janus, a foresight framework using multi-agent simulation to anticipate delayed operational risks from tool-using agents.
Research identifies 'contextual entrainment' in unimodal language models and proposes ENTRAP-VL to investigate its manifestation in vision-language models.
Researchers propose PoTRE, a framework using multiple AI agents for complex reasoning to address LLM struggles with long-horizon planning and error correction.
Research finds self-supervision, not clinical supervision, primarily drives representational convergence in medical foundation models, affecting their clinical usability.
Research explores PortLLM, a training-free, data-free method for adapting LLMs, highlighting its short-term temporal portability for LoRA patches.
Research finds vision-language models are sensitive to input modality order (image-first vs. question-first), impacting performance, and proposes a test-time training method to mitigate this.
Researchers propose LaSEr-Edit, a method for localized, span-level editing of LLM outputs to enforce safety and consistency constraints more reliably than brittle instruction-based control.
Research proposes AugAbEx, a hybrid method for abstractive and extractive legal summarization to address limitations of prior approaches in legal AI.
Research demonstrates that intermediate layers of LLMs and speech models predict brain responses to language, identifying abstraction as key to alignment.
Researchers propose Hibiki-Zero, a new method for simultaneous speech-to-speech translation that eliminates the need for word-level aligned data.
Researchers introduced 'vocabulary dropout' to prevent collapse in LLM co-evolution, enhancing curriculum diversity in self-play systems.
Research identifies 'self-preference bias' in rubric-based LLM evaluation where models favor outputs from themselves or their family, skewing benchmarks.
Research investigates how human-AI co-authorship and large language model (LLM) edits impact the detection of an author's native language (L1) traces in non-native writing.
Research introduces KoRe, a method to represent LLM knowledge externally, addressing opacity, debugging difficulties, and hallucination issues in parametric knowledge.
Researchers propose a graph-based RAG method to reduce hallucinations in complex question answering by improving context retrieval for LLMs.
Researchers propose BITEMBED, a low-bit framework for LLM-based text embeddings to improve efficiency and reduce storage and bandwidth overhead.
Researchers propose a meta-learning framework for aligning multilingual large language models by leveraging preference data across languages.
New research proposes an "in-the-flow" agentic system optimization for LLM planning and tool use, addressing scalability and generalization limits of monolithic policies.
Research finds verbatim conversational chunks outperform LLM-extracted structured artifacts for long-conversation memory recall in retrieval-rerank-reasoning pipelines.
Research explores prompt programming for LLM cultural bias and alignment in strategic decision-making and document engineering tasks.
Research challenges the common practice of using emotion embedding similarity (e.g., emotion2vec) to evaluate emotional expressiveness in speech generation models.
New research benchmark, CEO-Bench, evaluates AI agents on long-horizon, real-world tasks requiring complex skill orchestration and adaptation.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion