Scoring & evidence
LLM-as-judge
Using a language model to grade another system's output.
Also known as: model judge
Definition
An LLM judge applies instructions or a rubric to score free-form answers and complex deliverables that are difficult to grade deterministically.
Why it matters
It enables evaluation at scale but introduces judge bias, inconsistency and possible preference for related model families.
Related concepts
- Rubric
Explicit criteria describing what a good answer or deliverable must contain.
- Human baseline
Performance achieved by an appropriate group of people on the same task.
- Inter-judge agreement
How consistently two or more graders assess the same outputs.