RESEARCHMonitorNEXT 12 MONTHS
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
A reproducibility study evaluates whether lightweight MLP probes on LLM latent activations reliably detect harmful prompts across model families.
Open source