RESEARCHMonitorNEXT 12 MONTHS
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research reveals output discrepancies across AI access modes like APIs and chat interfaces, undermining evaluation validity.