RESEARCHInvestigateNEXT 12 MONTHS
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Academic research demonstrates that LLM safety benchmarks fail to accurately predict model behavior when prompts undergo minor surface-form variations.
Open source