The Sparsity Whisperer
Researchers propose a new LLM pruning technique that preserves how MLP layers separate similar inputs to maintain model performance.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose a new LLM pruning technique that preserves how MLP layers separate similar inputs to maintain model performance.
KReF introduces a training-free retrieval framework for long-term time-series forecasting and predictive uncertainty estimation.
Researchers introduce CubicQuant, a parametric non-uniform weight quantization method designed to accelerate 1-8 bit LLM inference on GPUs.
An academic paper proposes a mathematical framework to explain artificial neural network inference logic as sparse symbolic patterns.
Researchers introduced a new KV cache compression method that dynamically allocates cache resources across layers and context slots.
Researchers introduce Accounting Graph Transformer (AGT) to forecast 13 key financial KPIs simultaneously from short-history SME ledgers.
Researchers developed a pruning method to simplify reinforcement learning policies converted into decision trees, maintaining performance.
Researchers propose a network-curvature-based edge sparsification method to reduce compute costs in large-scale dynamic graph learning.
Researchers propose Conformal Fusion, a framework to maintain mathematically calibrated confidence when multimodal AI models encounter missing data feeds.
Researchers propose Fairis, an aggregation method defending collaborative machine learning against adversarial fairness poisoning attacks.
Researchers propose Cascade, an LLM serving framework that uses SLO-aware latency budgets to optimize inference cost and throughput.
Researchers demonstrate that pre-inference routing based on document difficulty can optimize the cost-performance trade-off in extraction.
Researchers introduce LivePlan, a framework to monitor and correct plan drift and repetitive errors in long-horizon programming agents.
Research shows LLM memory compression routinely drops epistemic qualifiers (e.g., 'allegedly', 'tentatively'), losing critical nuance.
Researchers proposed Agent Memory Distillation (AMD), a training-free framework that transfers structured memory from large to small LLM agents.
Researchers propose CoinRAG, a framework optimizing KV cache reuse in long-context RAG by focusing on fine-grained information nuggets.
Researchers propose O2CP, an optimization-based online conformal prediction framework to improve uncertainty quantification in multi-step time series forecasting.
Researchers propose a provable set-level inference framework to statistically identify whether specific datasets were used to train an LLM.
Research demonstrates that standard time-series benchmarking hides the fact that simple classical models often match or beat deep learning.
Moonshot AI released Kimi K2.5, an open-source multimodal agentic model featuring joint text-vision optimization and an Agent Swarm framework.
An arXiv research paper analyzes how 35 industry developers perceive, prioritize, and manage novel risks in agentic AI deployments.
Researchers introduce SkillTrace, a multi-trace provenance auditing tool designed to track the reuse and leakage of modular LLM-agent skills.
Research demonstrates standard top-k vector RAG fails on tabular financial reports, advocating for agentic, tabular parsing pipelines.
Researchers introduce H+ Embedding, a retrieval method balancing single-vector compression and expensive token-level late interaction.
Researchers introduce a topology-aware data movement framework to resolve networking bottlenecks in disaggregated LLM inference.
Researchers introduce TokTier, a stateful tokenization framework designed to eliminate redundant tokenization overhead in agentic LLM loops.
Researchers introduce FinanceHarness, an open framework designed to evaluate autonomous agentic workflows on financial deep research tasks.
Research explores whether AI agents can conduct open-ended AI research, introducing a new evaluation method beyond narrow tasks or peer review.
Researchers introduce Ratchet, an agentic lifecycle tool addressing the 0% performance gain of LLM-authored skills compared to human ones.
Research introduces CoSA, a new sparse attention mechanism designed to accelerate long-context inference in large language models by co-designing proxy and kernel.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion