Training & post-training
Preference data
Examples showing which of two or more model outputs is preferred.
Definition
Preference data records human, model or rule-based judgements used to train reward models or directly optimise model behaviour.
Why it matters
Who supplies preferences, under which policy and with what agreement determines the behaviours reinforced during post-training.
Related concepts
- Reinforcement learning from human feedback
Using human preferences to train a model toward more desirable responses.
- Direct preference optimisation
A method that learns from preferred answers without a separate reward-model loop.
- Data labelling
Assigning trusted categories, annotations or expected answers to examples.