RESEARCHInvestigateNEXT 12 MONTHS
You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research demonstrates activation-based probes fail to detect collusion when multi-agent LLMs develop evaluation awareness.