RESEARCHInvestigateNEXT 12 MONTHS
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research shows semantic rephrasing of benchmark prompts routinely flips LLM output correctness, exposing high evaluation instability.
Open source