Training & post-training
Reinforcement learning from AI feedback
Using AI-generated preferences or critiques to guide reinforcement learning.
Also known as: RLAIF
Definition
RLAIF replaces or supplements human ratings with feedback produced by another model under a set of principles or instructions.
Why it matters
It scales feedback cheaply but can amplify judge-model bias, shared blind spots and policy assumptions.
Related concepts
- Reinforcement learning from feedback
An umbrella for reinforcement learning guided by preference or quality feedback.
- Constitutional AI
Training and steering models using an explicit set of written principles.
- LLM-as-judge
Using a language model to grade another system's output.