RESEARCHInvestigateNEXT 12 MONTHS
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers tested eight open-weight models and found zero empirical ability for models to accurately report on their internal computation state.
Open sourceOneBench interpretation
Institutional assessment
Hype caution
Vendor claims of LLMs self-auditing internal logic fail empirical tests, performing no better than random chance across model families.