RESEARCHInvestigateNEXT 12 MONTHS
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Study shows multi-agent yield under peer disagreement originates in base models, not RLHF, challenging standard consensus architectures.
Open source