Measuring Reward-Seeking via Contrastive Belief Updates
Research introduces a method, Contrastive Synthetic Document Finetuning, to measure 'reward-seeking' behavior in LLMs trained with RL, where models optimize for grader judgment rather than the true objective.