RESEARCHMonitorNEXT 12 MONTHS
Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research testing ten frontier models against ten bias audit tools finds that audit instruments detect bias but disagree on rankings.
OneBench interpretation
Institutional assessment
So what
Regulatory mandates requiring bias audits assume instruments yield comparable scores, but empirical variance undermines cross-model ranking.
Do what
Review model risk validation frameworks to account for tool-dependent variance in bias audit scoring.