RESEARCHInvestigateNEXT 12 MONTHS
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
A study across 15 OpenAI models indicates safety alignment alters explicit gender bias into implicit forms rather than eliminating it.
OneBench interpretation
Institutional assessment
So what
Standard safety metrics and automated evaluation tools may mask underlying bias, creating unmitigated compliance risks under consumer protection standards.
Do what
Review model validation frameworks with the team responsible for AI governance to test for implicit bias beyond standard benchmark scores.