Toward Localizing and Repairing Bias in Transformer Attention Heads
Research explores localizing and repairing bias within specific transformer attention heads, moving beyond input-output or retraining methods.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research explores localizing and repairing bias within specific transformer attention heads, moving beyond input-output or retraining methods.
Researchers propose an agentic AI system for autonomous neural operator discovery using virtual laboratories that interact in a citation-based economy.
Research identifies 'Contextual Sycophancy' in reinforcement learning, where feedback sources are biased only in critical contexts, escaping standard robustness assumptions.
Research proposes a white-box instrument using hidden deterministic finite automata to assess if RL agents learn latent task states or shortcuts.
AgentLens introduces a new benchmark for coding agents, evaluating full trajectory performance beyond simple pass/fail metrics.
New research presents an exact algorithm for Data Shapley values in weighted k-nearest-neighbor regression and soft-label prediction.
New research proposes Hyperellipsoid Density Sampling, an efficient method to accelerate high-dimensional numerical optimization problems beyond traditional QMC.
Researchers propose a mobility-aware cache framework to improve the scalability and reduce computational cost of LLM-based human mobility simulations.
Research explores using spiking neural networks with active dendrites for energy-efficient simultaneous multi-task reinforcement learning in agents.
Research explores Stochastic Quantum Spiking Neural Networks with quantum memory and local learning, merging neuromorphic and quantum computing.
Research explores 'ontological inversion' where a predictive system's internal model permanently displaces its original environmental understanding.
New research proposes "significance-first splitting" for causal trees, aiming to improve heterogeneous treatment effect detection and inference validity.
Research introduces an audited protocol for detecting unique functional fingerprints in neural networks after convergence to a shared low-dimensional geometry.
Research introduces PluRel, a method for generating synthetic multi-table relational databases to train Relational Foundation Models (RFMs), overcoming privacy hurdles.
Research introduces a method, Prior-Aligned Training with Subset-based Attribution Constraints, to improve model reliability by aligning decisions with human priors.
LiteTopK is a new GPU kernel for sparse attention's Indexer-TopK operation, designed to reduce memory traffic and synchronization overhead for long-context LLMs.
Research identifies and proposes a method to mitigate 'shape-prior shortcutting' in single-shot fringe projection profilometry (FPP) networks.
Research challenges conventional LLM model merging, finding that expert training duration beyond optimal validation loss improves merged model quality.
Researchers propose BattVAE-GP, a hybrid physics-probabilistic framework for generative modeling of long-horizon battery degradation with uncertainty quantification.
Research introduces a new benchmark for institutional equity holdings prediction using temporal graph machine learning on SEC 13F filings.
Research characterizes communication and computation budget for decentralized gradient descent (DGD) to achieve target error levels.
Research finds striking zero-shot performance in multimodal benchmarks for computational pathology is compromised by data leakage at patient and institutional levels.
Research proposes SlimPer, a Transformer-based personalization model for recommendation systems that reduces tensor size and compute for efficiency.
MESH research proposes a unified retrieval model to handle diverse, heterogeneous content (fresh, long-tail) by addressing the Scaling Bias of Heterogeneity.
Research proposes a method to optimize synthetic data generated by diffusion models for improved few-shot medical image classification, focusing on usefulness for downstream tasks.
Research proposes a qubit-efficient quantum search for Hyperdimensional Computing (HDC) decomposition, addressing the computational challenge of scaling.
Mirror Theory introduces viable path entropy (VPE) as a measure of an intelligent system's capacity for coherent, verified continuations under reflection.
CARE-LoRA, a new research method, improves LoRA's memory efficiency for fine-tuning large pre-trained models by compressing activation reconstruction.
Research finds KV-cache compression methods perform differently under query-aware vs. query-agnostic (re-use) protocols, impacting real-world efficiency.
Research proposes a method to assess causal graph validity by falsifying candidate graphs based on their ability to explain outlier event propagation.
New research proposes the Environment Parameter Gradient Theorem for jointly optimizing reinforcement learning policies and environment design parameters.
Research explores methods for near-optimal learning of Gaussian Sobolev Operators, focusing on improving approximation guarantees and computational efficiency.
Research explores reproducible reservoir computing with thermally driven superparamagnets, aiming for robust performance under environmental conditions.
Research explores Sparse Autoencoders for improved out-of-distribution (OOD) detection by leveraging intermediate neural network layers.
Research introduces directional constraints for efficient exploration in Safe Reinforcement Learning, aiming to improve learning speed under safety constraints.
GenDiff, a new diffusion model, improves low-dose CT reconstruction by incorporating dose and anatomy awareness for better generalization across clinical settings.
Research introduces VQCSim, a compile-once statevector simulator addressing overheads in hybrid quantum-classical ML by optimizing static variational circuits.
Research introduces PFAdapter, a method for federated fine-tuning of Multimodal Large Language Models (MLLMs) using hierarchical LoRA decomposition.
Research uses machine learning, specifically XGBoost, to improve high-altitude Clear Air Turbulence (CAT) prediction in U.S. airspace.
Video diffusion models struggle with causal chains, degrading performance as interaction length increases in multi-ball hard-sphere dynamics.
Research indicates that spectral predictability scores for time-series do not accurately predict the utility of adding context (longer lookback, RAG).
Researchers introduced CoCo, a new loss function designed to create normalized, geometrically optimal embeddings for better classification and faster convergence.
Research proposes AVQ-Attention, an adaptive vector-quantized attention mechanism to improve transformer efficiency by optimizing codebook capacity.
Research characterizes features enabling effective 'grokking' in neural networks, showing how structured representations accelerate generalization.
New research proposes EG-VAR, a Lean 4-based architecture using tool-attestation and formal proofs to eliminate LLM hallucination in empirical inference.
AdaPCLA improves generative models for longitudinal EHR data by addressing underrepresentation of tail events, enhancing fidelity for rare subpopulations.
Research introduces Finite-Time Spectral Sensitivity (FTSS) as a gradient-free metric to diagnose flow geometry and memorization in continuous-time generative models.
New research proposes source-grounded feature inversion, a method to interpret neural network features by linking them directly to input samples.
Research introduces PoPE, a placebo-controlled evaluation methodology for measuring self-repair in frozen small code LLMs using execution counterexamples.
A new benchmark framework, OOD-RL-Bench, evaluates out-of-distribution (OOD) detection in reinforcement learning (RL) agents.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion