RESEARCHInvestigateNEXT 12 MONTHS
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers propose Self-Guided Adaptive Safety Alignment (SGASA) to help reasoning models synthesize and internalize safety guidelines.
Open source