RESEARCHMonitorNEXT 12 MONTHS
Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Researchers introduced an activation-space whitening method to detect LLM policy violations without retraining or high latency.
Independent coverage
OneBench grouped these reports as coverage of the same underlying development. Open each source to compare the evidence.