HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
HiKV proposes a novel algorithm-hardware co-design to compress the KV cache in LLM decoding, tackling memory bottlenecks for long-context models.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
HiKV proposes a novel algorithm-hardware co-design to compress the KV cache in LLM decoding, tackling memory bottlenecks for long-context models.
Research explores making large text-to-image diffusion models more interpretable and manipulable for creative uses, focusing on interactive explainability.
Research proposes a convex optimization framework to generate theoretical correlation matrices with graph-based sparsity patterns, improving matrix completion.
Research proposes TRACE-ROUTER, a new routing mechanism for agentic AI applications that optimizes LLM selection based on long-horizon, task-level outcomes.
Research proposes Generalized Gaussian Temporal Difference Error (GGD-TDE) for uncertainty-aware reinforcement learning, addressing non-Gaussian TD residuals.
Research proposes contract-based incentives for federated learning (FL) to prioritize high-quality data contributions during critical early training periods.
Research explores a Spatially-Enhanced Temporal Fusion Transformer for interpretable multi-output prediction in parametric dynamical systems with time-varying inputs.
Research explores meta-learning for speaker-dependent voice fatigue models to improve performance and efficiency over traditional mixed-effect models.
Research benchmarks federated learning strategies for in-hospital mortality prediction using heterogeneous and imbalanced clinical data.
Researchers introduced LiMuon, an optimizer designed for large model training, claiming reduced sample complexity and memory usage compared to prior Muon variants.
Research introduces SurvDiff, a diffusion model designed for generating synthetic survival data, addressing challenges of incomplete event information.
Research explores safety in In-Context Reinforcement Learning (ICRL), addressing unexamined test-time behavior for real-world deployments.
Research improves differentially private stochastic gradient descent (DP-SGD) accuracy by correlating privacy noise across iterations using model curvature.
Research identifies numerical fragility in Transformer models due to low-precision execution, proposing a layer-wise risk estimator and controller.
Research paper gp2Scale proposes a method for exact Gaussian Processes on up to 10 million data points, improving scalability without approximations.
Research proposes a Decentralized Multi-Agent Swarm (DMAS) architecture using autonomous agents for security in Industrial IoT (IIoT) environments.
Research introduces a supervisory runtime stability framework for neural network training to detect and recover from severe destabilizing updates.
Research proposes a layer-wise LoRA fine-tuning method using a similarity metric to improve LLM predictive performance efficiently.
Research introduces Heavy-Tailed Principal Component Analysis (HTPCA) for robust dimensionality reduction in heavy-tailed data and impulsive noise.
Research introduces Hierarchical Online Learning of Multiscale (HOLM) models, combining online latent-cause inference with hierarchical Bayesian models.
DriftXpress introduces an accelerated formulation for 'drifting models,' a new paradigm for one-step generative modeling that reduces inference costs.
Research explores short-term to long-term memory transfer for knowledge graphs in reinforcement learning under partial observability.
Research identifies that smaller LLMs inherently offer higher policy-level diversity for Group Relative Policy Optimization (GRPO), improving coherent trajectories.
Researchers propose Supervised Memory Training (SMT) to address gradient issues and parallelism limits in training recurrent neural networks (RNNs).
Research proposes a theory of indecisions for selective hypothesis testing to minimize abstention rates while maintaining target accuracy in high-risk scenarios.
Research on two-layer neural networks explores optimal generalization and learning transitions near the interpolation threshold for large models.
PCS-UQ introduces a framework for robust uncertainty quantification in high-stakes ML domains, integrating prediction-checks and bootstrapping.
Research paper explores statistical mechanics to better understand extensive-width Bayesian neural networks near interpolation, bridging theory-practice gap.
Research explores differentially private federated learning for imbalanced clinical data using SMOTETomek and FedProx to balance privacy and utility.
Research introduces Wasserstein Gradient Flows for scalable and regularized barycenter computation, improving aggregation of probability measures.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion