Risk & governance
Red teaming
Adversarial testing designed to uncover failures and unsafe behaviour.
Definition
Red teaming uses deliberate attacks, edge cases and misuse scenarios to probe how a system fails under pressure.
Why it matters
It exposes vulnerabilities that representative quality tests often miss, but cannot prove the absence of unknown attacks.
Related concepts
- Prompt injection
Instructions in untrusted content that attempt to redirect an AI system.
- Guardrail
A technical or procedural control that constrains system behaviour.
- Critical-error rate
The share of cases containing a failure with severe consequences.