RESEARCHInvestigateNEXT 12 MONTHS
Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Research identifies a critical security vulnerability in reasoning models where adversarial inputs starve safety monitors of token budget.
Open source