RESEARCHMonitorWATCHLIST
Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
New research shows extreme-value tail-index metrics used in LLM evaluation are fragile, challenging advanced risk-estimation frameworks.