Measuring the Depth of LLM Unlearning via Activation Patching
Research introduces activation patching to detect residual knowledge in LLMs that standard output-level unlearning metrics fail to catch.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces activation patching to detect residual knowledge in LLMs that standard output-level unlearning metrics fail to catch.
Researchers published PropUQ-MAS, a framework to quantify error propagation and uncertainty compounding across multi-agent LLM workflows.
ExecRubrics proposes executable code-based rubrics to replace ambiguous natural-language LLM judges for long-form evaluations.
Researchers released Thinkingbox, a sandbox benchmark evaluating AI agents on stateful business workflows and multi-turn tool reliability.
Researchers introduced a self-diagnosing framework that identifies root causes of AI failures under distribution shift beyond basic OOD detection.
New T2MO framework optimizes enterprise LLM coding assistant routing by accounting for retries, escalations, and developer wait times.
Researchers introduced PuMVR, a Punjabi multimodal benchmark revealing that state-of-the-art vision-language models fail at multi-script decoding.
Academic research demonstrates that LLM safety benchmarks fail to accurately predict model behavior when prompts undergo minor surface-form variations.
Researchers propose conformal prediction methods for LLMs that maintain mathematical reliability guarantees even when prompt configurations shift.
Research shows LLM sycophancy—abandoning correct answers under user pressure—is a conversational dynamic rather than a fixed model trait.
Researchers introduced ArabicDialectSafety, a human-curated safety dataset of over 25,000 prompts across six Arabic dialects.
Researchers introduced DRIP-R, a benchmark designed to evaluate how LLM agents reason and make decisions under ambiguous domain policies.
Researchers propose DuplexGen, a method for synthesizing adaptive, scenario-specific human-AI turn-taking behaviors for voice dialogues.
PeopleSearchBench is an open-source benchmark for evaluating AI-powered people search platforms across recruiting, sales prospecting, and expert search.
SyRuP, a new method, aims to improve LLM adherence to complex system prompts during decoding without requiring model tuning or reranking.
Researchers propose "masked distillation" to internalize Chain-of-Thought reasoning in language models, reducing inference latency and cost.
Research tests sparse attention mechanisms, finding attention patterns do not reliably indicate which parts of context are used by large language models for answers.
Research highlights lack of systematic justification in critical design decisions for universal multilingual Named Entity Recognition models.
Research finds sparse autoencoder (SAE) features' causal roles vary across SAE families and layer depths, challenging their stability for LLM interpretation.
Research finds deterministic KV-cache eviction prevents consistent error estimation, while randomized eviction restores identifiability for attention-output error.
Research identifies limitations of cosine similarity in learned representations, particularly in handling radial variation and anisotropy.
Research introduces a structure-preserving Physics-Informed Neural Network (PINN) for the Korteweg–de Vries (KdV) equation, enhancing long-term stability.
Researchers introduced a quantum graph convolutional architecture for unsupervised learning designed for noisy intermediate-scale quantum (NISQ) hardware.
Research explores if Large Language Models (LLMs) can benefit from “experiential abstractions,” similar to humans distilling experience into reusable strategies.
Researchers introduced Spaghetti Architect, a tool that generates synthetic, labeled code datasets to address issues with uncontrolled mined corpora.
New research introduces Off-Context GRPO, an RL method that provides privileged guidance during training to overcome zero-reward plateaus in LLM reasoning.
Research explores using foundation model embeddings for interpretable, region-based brain MRI classification, improving anatomical interpretability.
Research explores optimal depth for graph neural networks (GNNs) on sparse graphs for node classification, focusing on message-passing limits.
Research introduces SALT, a Salience-Aware Lexical Trie method for long-context compression addressing 'theme collapse' in LLM inference.
WorldCupArena is a new dynamic benchmark on arXiv for evaluating LLMs and deep-research agents on real-time football match forecasting.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion