RESEARCHInvestigateNEXT 12 MONTHS
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
New research argues that current methods for evaluating LLM activation explanations are structurally flawed, failing to penalize individual false claims.
Open source