DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
DORA is an asynchronous reinforcement learning system that speeds up LLM post-training by overlapping generation with model training, addressing the rollout phase bottleneck.