RESEARCHInvestigateNEXT 12 MONTHS
Safety Training May Persist Through Helpfulness Optimization in LLM Agents
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
ArXiv research demonstrates safety training via DPO can persist during helpfulness optimization in multi-step agentic tool-use environments.
Open source