RESEARCHMonitorWATCHLIST
Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research shows binary-choice LLM benchmarks like TruthfulQA suffer from surface-level feature leakage that inflates accuracy scores.