RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
Factual evidence
What the source reports
Researchers introduced RAG-Stress, a protocol measuring how misleading retrieved evidence causes LLMs to override correct answers.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.CL — Computation and Language · RESEARCH
- Published by source
- 9 October 2026
- Collected by OneBench
- 10 Oct 2026, 03:01 UK
Stored source excerpt
arXiv:2610.11183v1 Announce Type: new Abstract: Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.
OneBench interpretation
Institutional assessment
So what
Standard RAG architectures in document search can actively degrade accuracy by forcing models to accept incorrect retrieved context.
Do what
Review evaluation protocols with the team responsible for model risk management when testing RAG applications on regulated data.