CausalGate: Causal Importance Distillation for Transformer Module Pruning
CausalGate introduces an intervention-guided framework for transformer module pruning, moving beyond observational heuristics to capture non-linear structural computations.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
CausalGate introduces an intervention-guided framework for transformer module pruning, moving beyond observational heuristics to capture non-linear structural computations.
Research shows coding agent token consumption varies significantly across programming languages (Python, Java, Rust, OCaml), impacting inference cost.
Research proposes a logit biasing method for adapting black-box language models for domain-specific tasks and privacy without full fine-tuning.
DS@GT submitted a language-routed RAG system to FinMMEval 2026 Task 1 for multilingual financial exam Q&A, covering English, Spanish, Greek, Chinese, and Hindi.
Research introduces AssumptionMiner, a method to extract, trace, and revise implicit assumptions made by LLMs during code generation from incomplete prompts.
Researchers propose Verbalized Particle Posterior (VPP) for Bayesian inference over natural language hypotheses in LLM-based models, addressing uncertainty and stability.
ConsistencyGate proposes a self-consistency admission control mechanism to prevent LLM agent memory contamination from hallucinated facts in external memory stores.
Research identifies a fundamental limitation in dual-encoder vision-language models where compositional queries like "A and not B" fail due to a Bag-of-Concepts effect.
Research on Multi-head Latent Attention (MLA) in DeepSeek-V2 details how it compresses key-value pairs, reducing KV-cache by 81% during inference.
Research explores hybrid retrieval for dynamic product ads, combining LLMs with traditional methods to balance semantic intent and cost.
AgentOmnia presents a framework for scaling LLM agents across To-Consumer, To-Business, and To-Employee applications through coordinated development.
Research reveals score-conditioned In-Context Learning (ICL) in LLMs corresponds structurally to policy gradient optimization, explaining iterative output improvements.
Research proposes a two-stage incremental framework for LLM-based dialogue systems to reduce response delays by generating prefatory responses early.
New research introduces a formal model to study hallucination in language generation, contrasting theoretical limits with practical LLM behavior.
Research proposes a new method, Weighted Banzhaf Interactions with Tree-Gram Parsing, to explain BiomedCLIP's medical task performance by addressing tokenizer fragmentation.
Omni-Prune proposes query-aware token pruning for omnimodal LLMs to reduce inference latency and GPU memory for audio-video inputs.
Research proposes a method to audit LLM alignment controllability, measuring how far responses can be steered from a 'resting point' via system prompts.
Researchers developed HiTMS, a high-throughput multi-stream linguistic steganography framework to conceal secret data in LLM responses.
Research highlights unaddressed security and governance risks of third-party API routers in agentic AI workflows, especially regarding data inspection and modification.
Latent-LoRA introduces a gradient-free routing mechanism for LoRA adapters to prevent catastrophic forgetting in continual learning tasks for LLMs.
Research finds AI models GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, and Grok 4.5 show bias in naming individuals based on citation type.
Research introduces 'reality monitoring' for LLMs, showing models struggle to distinguish their own output from user input, impacting accuracy.
Research explores machine unlearning by examining mode connectivity to improve understanding of loss landscape and optimization geometry.
New research proposes Adaptive Control of Training-Inference Discrepancy (ACRL) to stabilize RL training for LLMs by mitigating issues from architectural separation and precision differences.
Research identifies a gap in conjunctive cross-page retrieval, where current systems struggle to confirm all evidence for multi-part requests within a document.
Research proposes Gubernaut, a model-agnostic runtime control layer, to manage LLM agent "propensity failures" like escalation, sycophancy, and perseveration.
Research proposes role-stratified conformal risk control for LLM tool calls to prevent high-risk argument failures from being masked by benign ones.
Research evaluates how model scale and quantization affect uncertainty signals in Vision-Language Models (VLMs) under image degradation, impacting confidence for deferral decisions.
Research reveals reward models (RMs) misallocate memorization to easy preferences, memorize dataset-specific shortcuts, and overgeneralize heuristics.
Research explores sparse autoencoders (SAEs) to better link SAE features to LLM behavior, addressing inconsistencies in causal effects and steering.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion