RESEARCHInvestigateNEXT 12 MONTHS
Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research demonstrates Best-of-N inference scaling causes 'safety hacking' when imperfect safety filters are paired with reward maximization.
Open sourceOneBench interpretation
Institutional assessment
So what
Inference-time guardrails relying on proxy safety models and reward scaling create systemic failure modes that allow unsafe outputs to bypass enterprise controls.
Do what
Ask your model risk team to audit post-processing safety filters for compound optimization failures in production inference pipelines.