AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
AngelSpec proposes a new speculative decoding method combining multi-token prediction and block-parallel diffusion for LLM inference acceleration.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
AngelSpec proposes a new speculative decoding method combining multi-token prediction and block-parallel diffusion for LLM inference acceleration.
Research paper proposes evaluating multimodal LLMs for clinical diagnostics based on multi-turn interactions and progressive information disclosure, reflecting real-world clinical practice.
Researchers introduced Polistemics, a new benchmark to evaluate Large Language Models (LLMs) as responsible mediators of political information in elections.
Research identifies methods to detect knowledge inconsistencies across multimodal data (text, tables, knowledge graphs) from sources like Wikipedia and Wikidata.
Research finds instruction-tuned models reuse human syntax more than humans themselves, indicating a different form of linguistic adaptation.
Research paper introduces UniMem, a hybrid memory architecture for LLM agents to improve performance on continuous, evolving task streams by combining episodic and parametric memory.
Research identifies 'prefix failure' in on-policy distillation, where student LLMs commit to wrong reasoning paths, and proposes 'trajectory-relayed distillation' to address it.
Research explores multi-objective structured pruning techniques for LLMs to optimize for lower latency and smaller model size, crucial for edge deployment.
DocAnnot, a new framework using Large Vision Language Models and a novel Spatially Informed Contextual Matching algorithm, accelerates Key Information Extraction (KIE) dataset creation.
VLD-RAG is a research paper exploring agentic multimodal retrieval-augmented generation for question answering over long, visually-rich multi-page documents.
Research explores treating language as a material for creative LLM interaction beyond directive prompting, focusing on open-ended and associative uses.
Research explores how text chunk size in Retrieval-Augmented Generation (RAG) systems impacts Large Language Model (LLM) performance and generation quality.
Research finds LLMs trained on less diverse language data exhibit increased in-context 'scheming' or misaligned objective pursuit.
Research explores LLMs for specialized terminology translation, evaluating their ability to find equivalents for English to French compared to traditional corpora.
Researchers propose GLIDE, a hybrid attention mechanism combining sliding-window and linear aggregation to reduce KV cache overhead for long-context LLM inference.
An arXiv paper details Ontario Power Generation's evolving RAG pipeline, from naive RAG to deep agentic retrieval, for regulatory compliance.
Research suggests retrieval failures, not hallucinations, will be the primary limiting factor for LLM-based clinical AI tools, shifting focus to recall errors.
Research paper explores how Large Language Models internally collapse reading and writing into a single entangled autoregressive process, unlike human brains.
TraceBound is a diagnostic protocol for adaptive knowledge-graph retrieval systems, studying how pre-retrieval 'thinking' affects robustness.
Research investigates prompt engineering's influence on Small Language Models (SLMs) for guarded query routing, evaluating 22 models on GQR-Bench.
SourceMinds at CheckThat! 2026 presents a multi-agent pipeline for fact-checking article generation, incorporating NLI-grounded citation auditing.
Research quantifies RAG's contribution to safety in LLMs for mental health, enhancing intent detection and mitigating hallucination.
Research reveals LLM-based recommender systems are vulnerable to position bias, enabling attackers to promote items by reordering candidates.
Agent Retrieval Bench is a new benchmark for evaluating the context acquisition stage of coding agents, focusing on file-level retrieval in repositories.
Research explores uncertainty in retrieval-augmented generation for code, focusing on relevance, compatibility, and completeness of retrieved information.
Mage-VL is a new research model focusing on efficient, real-time streaming multimodal understanding by selectively encoding dynamic visual information.
Researchers propose Addressable Recall Compaction (ARC), a new context-management framework for LLM agents to overcome context window limits by separating archival storage from active context.
Research systematically investigates instability in reinforcement learning for Small Language Models (SLMs) in the 70-500M parameter range.
Research explores emotional contagion in multi-agent LLM simulations, where agents perceive, appraise, and propagate affect based on personality and context.
CLBench-V introduces a new benchmark for evaluating multimodal context learning, focusing on models' ability to learn from diverse contexts beyond text.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion