Search signals, briefings, benchmarks and glossary terms.
Catalogue reviewed 26 Jul 2026·32 evaluation assets across 10 bank domains·2 known public-coverage gaps·Last ingest 31 Jul, 20:36 UK
The OneBench evaluation position
A leaderboard is a screening tool, not a model approval.
What a bank needs
Combine external benchmarks with a versioned use-case golden set, adversarial control tests, human baselines, operational fitness and continuous evaluation against the exact system being deployed.
Evaluation coverage matrix
Start here: read across a bank function to see where reusable evaluation evidence exists—and where public benchmarks stop. Every cell opens the catalogue filtered to that function and evaluation layer.
Amber always requires institution-owned cases
Evaluation asset coverage by bank function and evaluation-stack layer
Bank function
External screenPublic benchmarks and standards used to shortlist systems.
Golden setRepresentative institution-owned cases for the intended use.
ControlsSafety, conduct, security and policy-boundary testing.
OperationalWorkflow, tool-use, cost, latency and human-escalation evidence.
ContinuousRepeatable regression and production-change evaluation.
No mapped asset Some coverage Deeper coverage Local golden set required
Cybersecurity: thin by design
This catalogue covers AI-specific security evaluation. General cyber control and penetration-testing suites remain outside its scope.
Software engineering: thin by design
Coverage is limited to evaluation of AI coding systems, not conventional software-development lifecycle test suites.
Benchmarks and evals 101
Learn the concepts in four minutes: what a benchmark is, what an eval adds, how evidence moves from market comparison to bank approval, and how to challenge a headline score.
Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.
Primary sources linked
13 matching evaluation assets0 selected for comparison
Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.
OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.
Keep pace with finance AI evaluation
The daily briefing flags material benchmark and assurance developments. For catalogue governance or team access, speak to OneBench directly.
Free. Daily at 06:30 UK. Unsubscribe with one click.