RESEARCHInvestigateNEXT 12 MONTHS
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research explores 'input-only suppression' of LLM evaluation-awareness latents to prevent models from detecting safety evaluations.
Open source