RESEARCHMonitorNEXT 12 MONTHS
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers identify 'word recovery' as the mechanism enabling LLMs to process and reconstruct noisy or character-level tokenized inputs.
Open source