RESEARCHInvestigateNEXT 12 MONTHS
Measuring the Depth of LLM Unlearning via Activation Patching
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research introduces activation patching to detect residual knowledge in LLMs that standard output-level unlearning metrics fail to catch.
Independent coverage
OneBench grouped these reports as coverage of the same underlying development. Open each source to compare the evidence.