RESEARCHInvestigateNOW
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Audit of 52,988 API calls shows closed-source LLM judges fail basic temporal stability tests on identical prompts over time.