QR-Erase: Efficient Subspace-Based Machine Unlearning with Layer Localization
Researchers introduce QR-Erase, a machine unlearning method using Pivoted QR decomposition to remove targeted data without full retraining.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduce QR-Erase, a machine unlearning method using Pivoted QR decomposition to remove targeted data without full retraining.
Researchers introduce Population Aligned Language Models (PALMs) to simulate country-specific cultural values and population preferences.
Researchers introduce DocNavRAG, a framework combining GraphRAG and agentic navigation to query complex, structured document collections.
An academic paper proposes evaluating language models by analyzing semantic structures in embedding spaces rather than behavioral outputs.
Researchers analyzed LLM evaluations of non-native Japanese writing, finding models mimic human biases regarding fluency, status, and solidarity.
Researchers propose RING, a framework replacing external RAG retrievers with internal Mixture-of-Memory Experts trained via RL.
Researchers show that KV cache compression in reasoning LLMs can preserve final-answer accuracy while degrading supporting rationales.
Researchers prove that Block Sparse Flash Attention optimizations silently alter model outputs, introducing a counterfactual audit framework.
Researchers introduce RADAR, a diagnostic framework to detect behaviorally coupled and redundant criteria in rubric-based LLM-as-judge pipelines.
Researchers propose CRISP, a method to train search agents that prune redundant queries and irrelevant web observations to lower compute costs.
Researchers propose TRAM, an auxiliary memory architecture to prevent reasoning loss in Multimodal Large Reasoning Models over long context trajectories.
Researchers propose a method to validate teacher guidance trajectories in on-policy distillation, preventing compounding errors in agentic LLMs.
Researchers propose IACM-RL, a framework using context management and reinforcement learning to stabilize agentic tool invocation.
Researchers propose a method for LLMs to self-improve by internalizing transient interaction experiences into model parameters.
Researchers propose 'Region Grafting' to optimize failed agentic workflow segments at inference time using execution signals without complete re-runs.
Researchers introduced ScrambleToolBench, a benchmark designed to evaluate how autonomous agents discover and use undocumented tools.
Researchers released PredAct-Bench, a benchmark evaluating agentic LLM performance when interacting with noisy external APIs and tools.
A systematic study evaluates training-free versus training-based intent classification methods for routing prompts to specialized LLMs.
Researchers propose CTRAG, an in-context retrieval-augmented framework designed to automate regulatory compliance checking using LLMs.
Researchers prove that training LLMs to abstain from answering using standard RLHF with KL-divergence penalties can cause training collapse.
Researchers demonstrate that AI agents draw different conclusions from identical numerical data based on changes to the scenario's framing.
A paper analyzes GraphRAG vs vector RAG performance across varying embedders, corpora, and judges, finding verdicts depend on citation placement.
Researchers find that LLM persona panels can easily pass broad statistical validation tests while failing to replicate granular human behaviors.
Researchers introduce a dataset of 1,600 human-written hallucination samples to benchmark vision-and-language models independently.
Wix researchers published a three-stage pipeline for LLM agents that filters tool selection based on real-time user account state eligibility.
Researchers introduce CompressAgent to evaluate how prompt compression affects the operational reliability of tool-using LLM agents.
Researchers propose a self-evolving curriculum for reinforcement fine-tuning to overcome data scarcity in complex reasoning tasks.
Researchers introduce V-Mem, a modality-routed retrieval architecture designed to improve visual query accuracy in multimodal agents.
An academic study finds that adversarial self-play does not improve legal reasoning in models, yielding a controlled negative result.
Researchers introduced ConfBench, a calibration-specific benchmark to evaluate if vision-language models' confidence scores are reliable for document extraction tasks.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion