RESEARCHMonitorNEXT 12 MONTHS
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research auditing 256 open-weight chat tokenizers reveals all are vulnerable to control-token forgery via user-controlled text.
OneBench interpretation
Institutional assessment
So what
Control-token forgery exposes open-weight chat models in enterprise environments to prompt injection and role-spoofing attacks.
Do what
Review input validation and tokenization sanitisation controls with the team responsible for internal model security.