RESEARCHInvestigateNOW
Can LLMs Reliably Self-Report Adversarial Prefills, and How?
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Academic study of 10 open-weight models shows they cannot reliably self-detect when they have fallen victim to adversarial prefill attacks.
Open source