FRONTIER LABMonitorWATCHLIST
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Apple ML Research
Factual evidence
What the source reports
Apple ML Research proposes DACA-GRPO, a denoising-aware credit assignment method for reinforcement learning in diffusion language models.