RESEARCHInvestigateNEXT 12 MONTHS
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers propose a zero-call protocol-level identifiability audit to verify whether LLM reasoning benchmarks measure intended properties.
Open source