A leaderboard is a screening tool, not a model approval.
Use public benchmarks to screen models; approve deployments with bank-owned use-case, control and operational tests.
Use public benchmarks to screen models; approve deployments with bank-owned use-case, control and operational tests.
Start here: read across a bank function to see where reusable evaluation evidence exists—and where public benchmarks stop. Every cell opens the catalogue filtered to that function and evaluation layer.
Amber always requires institution-owned cases| Bank function | External screenPublic benchmarks and standards used to shortlist systems. | Golden setRepresentative institution-owned cases for the intended use. | ControlsSafety, conduct, security and policy-boundary testing. | OperationalWorkflow, tool-use, cost, latency and human-escalation evidence. | ContinuousRepeatable regression and production-change evaluation. |
|---|---|---|---|---|---|
| Finance & investment | 13assets | 0assetsLocal set required | 0assets | 8assets | 0assets |
| Legal & compliance | 7assets |
Learn the concepts in four minutes: what a benchmark is, what an eval adds, how evidence moves from market comparison to bank approval, and how to challenge a headline score.
Open the full glossary →A benchmark is a standardised test: the same tasks, conditions and scoring method are applied to multiple models or systems.
Use public benchmarks to shortlist models and expose obvious weaknesses.
Understand this layer →02Use-case golden setTest representative, difficult and high-risk cases drawn from the bank's work.
Understand this layer →03Controls & adversarialProbe leakage, entitlements, prompt injection, conduct, bias and unsafe actions.
Understand this layer →04Operational fitnessMeasure latency, cost, availability, fallbacks, auditability and human escalation.
Understand this layer →05Continuous evaluationRerun on model, prompt, data, tool and policy changes; monitor production drift.
Understand this layer →Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.
Primary sources linked| Compare | |||||||
|---|---|---|---|---|---|---|---|
| Public benchmark | Finance & investment, Legal & compliance, Operations & agents | Emerging | Open dataset | 2026-07-26 | Mercor Open ↗ | ||
| Internal test pack | Marketing & customer, Legal & compliance | Build internally | Internal | 2026-07-25 | Institution-owned Open ↗ | ||
| Internal test pack | Procurement & third party, Legal & compliance, Operations & agents | Build internally | Internal | 2026-07-25 | Institution-owned Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open source | 2026-07-26 | Handshake AI Research and collaborators Open ↗ | ||
| Public benchmark | Marketing & customer, Operations & agents | Established | Open dataset | 2026-07-26 | PolyAI Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents, Knowledge & RAG | Emerging | Open source | 2026-07-26 | Rogo Technologies Open ↗ | ||
| Public benchmark | Legal & compliance, Knowledge & RAG | Growing | Public methodology | 2026-07-26 | Harvey Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open dataset | 2026-07-26 | HiThink Research Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open source | 2026-07-26 | Longitude Labs and collaborators Open ↗ | ||
| Public benchmark | Legal & compliance, Procurement & third party | Established | Open dataset | 2026-07-25 | Stanford NLP Open ↗ | ||
| Public benchmark | Knowledge & RAG | Growing | Open source | 2026-07-25 | Meta and collaborators Open ↗ | ||
| Public benchmark | Marketing & customer, Operations & agents | Growing | Public methodology | 2026-07-25 | Salesforce AI Research Open ↗ | ||
| Public benchmark | Legal & compliance, Procurement & third party | Established | Open dataset | 2026-07-25 | The Atticus Project Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents, Knowledge & RAG | Growing | Open source | 2026-07-26 | Vals AI Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG | Established | Open dataset | 2026-07-25 | Patronus AI Open ↗ | ||
| Public benchmark | Finance & investment | Growing | Open dataset | 2026-07-26 | FinanceReasoning research team Open ↗ | ||
| Public benchmark | Finance & investment | Growing | Open source | 2026-07-25 | The FinAI Open ↗ | ||
| Public benchmark | Finance & investment | Established | Open dataset | 2026-07-25 | IBM Research and collaborators Open ↗ | ||
| Public benchmark | Legal & compliance, Operations & agents, Knowledge & RAG | Emerging | Open source | 2026-07-26 | Harvey Open ↗ | ||
| Public benchmark | Legal & compliance | Established | Open dataset | 2026-07-25 | Stanford and legal contributors Open ↗ | ||
| Public benchmark | Legal & compliance | Growing | Open dataset | 2026-07-26 | LEXam research consortium Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG, Operations & agents | Emerging | Open source | 2026-07-26 | Databricks Open ↗ | ||
| Public benchmark | Knowledge & RAG | Growing | Open dataset | 2026-07-25 | Galileo Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Growing | Open source | 2026-07-26 | SpreadsheetBench contributors Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG | Established | Open dataset | 2026-07-25 | NExT++ Open ↗ | ||
| Public benchmark | Operations & agents, Procurement & third party | Growing | Open source | 2026-07-25 | ServiceNow Research Open ↗ | ||
| Public benchmark | Operations & agents, Marketing & customer | Growing | Open source | 2026-07-25 | Sierra Research Open ↗ |
Patronus AI
Open-book financial question answering over public-company filings with evidence-linked answers.
Test research assistants, filing analysis, credit research and finance RAG before adding institution-specific documents.
Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.
Updated with the main intelligence pipelineMeta released its Muse Code AI coding agent in beta amid criticism for obscuring benchmark comparisons with OpenAI's Sol.
Open source ↗OpenAI has updated ChatGPT with its improved GPT-5.6 Sol model and expanded free user access to GPT-5.6 Luna.
Open source ↗Researchers propose Elbow-Based MoE Routing, a training-free inference plugin that dynamically selects Mixture-of-Experts active counts.
Open source ↗Researchers propose 'oblivious audits' to prevent model providers from detecting and manipulating regulatory or compliance evaluations.
Open source ↗Case studies from AI Snake Oil show agentic workflows fail at autonomous, open-ended research tasks despite vendor benchmarking claims.
Open source ↗Researchers evaluated prompt-based and trained methods to shorten model reasoning traces, measuring impacts on accuracy across GPQA and MMLU-Pro.
OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.
The daily briefing flags material benchmark and assurance developments. For catalogue governance or team access, speak to OneBench directly.
Free. Daily at 06:30 UK. Unsubscribe with one click.
| 2assets |
| 4assets |
| 0assets |
| Marketing & customer | 3assets | 1assetLocal set required | 1asset | 2assets | 0assets |
|---|
| Procurement & third party | 3assets | 1assetLocal set required | 1asset | 2assets | 0assets |
|---|
| Operations & agents | 13assets | 1assetLocal set required | 1asset | 13assets | 0assets |
|---|
| Knowledge & RAG | 9assets | 0assetsLocal set required | 0assets | 5assets | 0assets |
|---|
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion