Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
Researchers propose a functional reconstruction method for MLA draft models in speculative decoding to reduce KV cache memory overhead.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose a functional reconstruction method for MLA draft models in speculative decoding to reduce KV cache memory overhead.
Researchers introduce RLPF, a reinforcement learning method training code generation models to prefer faster, more efficient implementations.
A study reveals that 4-bit post-training quantization, while looking lossless on standard metrics, causes severe failures in multi-turn tool-calling LLM agents.
Researchers model LLM capability emergence as a rate-limiting nucleation process, explaining sudden capability jumps and plasticity loss.
A research paper analyzes why agentic AI performance degrades over long-horizon tasks, introducing trajectory-induced degradation concepts.
Researchers introduced EvoCause, a framework using LLMs to iteratively refine causal graphs for root cause analysis in complex IT infrastructures.
Researchers introduced CoT-Mediate, a framework evaluating if medical VLMs' reasoning causally drives predictions or merely mimics user expectations.
Academic research identifies mathematical failure modes in bilinear contrastive critics used for LLM reranking and best-of-K selection.
Researchers propose model compression techniques to reduce the high memory footprint of tabular foundation models like TabPFN.
Researchers propose Policy Gradient Steering (PGS), applying reinforcement learning to activation steering for dynamic LLM behavioral control.
Researchers propose Compliance2LoRA, a hypernetwork generating on-demand LoRA adapters to align LLMs to dynamic subsets of safety policies.
Researchers propose a Kalman-filter-based curriculum learning approach to optimize prompt difficulty selection during RL fine-tuning.
Researchers proposed a zero-shot transfer protocol for training GNNs on small graph replicas and deploying them on large graphs without retraining.
Researchers propose an expand-then-compress RL framework, arguing single-run RL models are biased teachers that fail to capture all valid reasoning paths.
A comparative study evaluates Large Language Models against LSTMs and tabular models for predictive process monitoring of event logs.
Researchers propose Contrastive Concept Importance to explain pairwise class decisions using automatically extracted concept representations.
Researchers evaluate reward design for Reinforcement Learning with Verifiable Rewards to improve LLM machine unlearning efficiency.
Researchers introduce TAPO, a post-training method optimizing LLM agents using dense environmental feedback rather than sparse task rewards.
Researchers introduce ClawTrack, a dual-assessment benchmark designed to evaluate both the task outcomes and trace-level reasoning of autonomous agents.
Research shows mathematically equivalent expert aggregation orders in sparse MoE models like DeepSeek-V4-Flash cause divergent outputs.
Researchers propose an Information Bottleneck approach to improve the mathematical faithfulness of explanations in time series forecasting.
Researchers propose LEDGERMIND, a framework for evaluating multimodal agents using a structured evidence ledger for step-by-step verification.
Researchers propose Adaptive Anticipatory Policy Trees to eliminate autoregressive decoding delays in multimodal computer-use agents.
Research trains a chain-of-thought (CoT) reasoning-enabled LLM to classify cybersecurity detections, aiming to reduce alert fatigue in SOCs.
Research explores integrating causality into algorithmic recourse to ensure recommended changes genuinely improve qualifications, not just game classifiers.
Research identifies a structural issue in on-policy self-distillation (OPSD) for LLMs, proposing $\beta$-OPSD as a more stable policy optimization approach.
KAISEN proposes a five-phase, reproducible audit pipeline for clinical risk models to identify and diagnose error rate disparities across patient subgroups.
Researchers demonstrate that Earth observation foundation models improve regional forest biomass monitoring, addressing data sparsity in carbon tracking.
Researchers introduce Divergence Decoding, a training-free inference framework that fuses generalist reasoning and specialist domain models.
Academic paper proves that retraining neural networks on pooled datasets can cause unexpected decision reversals, undermining model reliability.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion