RESEARCHInvestigateNEXT 12 MONTHS
Activation Probes Surface Code-Security Signals that the Model's Output Misses
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Probing activations of open-weight reviewer models surfaces hidden code vulnerabilities missed by the model's textual outputs.
Open source