Training & post-training
Synthetic data
Artificially generated examples used instead of or alongside observed data.
Definition
Synthetic data is produced by simulations, rules or models to expand, balance, protect or stress a training or evaluation dataset.
Why it matters
It can address scarcity and privacy but may reproduce generator errors, miss real-world tails or create misleading diversity.
Related concepts
- Data labelling
Assigning trusted categories, annotations or expected answers to examples.
- Data provenance
The origin, ownership and transformation history of data.
- Golden set
A curated set of representative inputs and trusted expected outcomes.