Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
Factual evidence
What the source reports
Researchers introduced the ASRD dataset to evaluate open-weight LLM safety robustness against non-canonical text inputs like emojis.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.CL — Computation and Language · RESEARCH
- Published by source
- 8 October 2026
- Collected by OneBench
- 9 Oct 2026, 03:01 UK
Stored source excerpt
arXiv:2610.09033v1 Announce Type: new Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.
OneBench interpretation
Institutional assessment
So what
Non-canonical text variations can easily bypass standard open-weight LLM safety guardrails, exposing deployed enterprise applications to unexpected risk.
Do what
Review input sanitization and safety testing protocols for open-weight models with the team responsible for model risk management.