RESEARCHMonitorWATCHLIST
Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers test Gemma-2 and Llama-3.1 agents trained via moral-reward RL, finding adversarial persona attacks break moral alignment.