RESEARCHInvestigateNEXT 12 MONTHS
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers introduced Groundedness Drift, a method using explanations to detect backdoors in black-box LLM classifiers.
Open sourceOneBench interpretation
Institutional assessment
So what
Backdoored classifiers used in automated content routing or sentiment analysis can be audited without knowing the underlying trigger phrase.
Do what
Ask your model risk validation team to assess Groundedness Drift techniques for third-party black-box classifier audits.