RESEARCHMonitorNEXT 12 MONTHS
Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research demonstrates a novel white-box attack using knowledge editing mechanisms to force unsafe output probabilities in LLMs.
Open source