RESEARCHInvestigateNEXT 12 MONTHS
How Context Attribution Handles What the Model Already Knows
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
New research identifies a critical flaw in LLM context attribution methods: they fail to distinguish between information retrieved from input context and knowledge encoded in model weights, producing unreliable scores.
Open sourceOneBench interpretation
Institutional assessment
So what
Attribution failures when context overlaps training data undermine explainability claims for RAG-augmented LLMs, complicating model risk assessments.
Do what
Brief your model risk team on the 'in-weight vs in-context' problem in current explainability tooling.