A leaderboard is a screening tool, not a model approval.
Use public benchmarks to screen models; approve deployments with bank-owned use-case, control and operational tests.
Use public benchmarks to screen models; approve deployments with bank-owned use-case, control and operational tests.
Start here: read across a bank function to see where reusable evaluation evidence exists—and where public benchmarks stop. Every cell opens the catalogue filtered to that function and evaluation layer.
Amber always requires institution-owned cases| Bank function | External screenPublic benchmarks and standards used to shortlist systems. | Golden setRepresentative institution-owned cases for the intended use. | ControlsSafety, conduct, security and policy-boundary testing. | OperationalWorkflow, tool-use, cost, latency and human-escalation evidence. | ContinuousRepeatable regression and production-change evaluation. |
|---|---|---|---|---|---|
| Finance & investment | 13assets | 0assetsLocal set required | 0assets | 8assets | 0assets |
| Legal & compliance | 7assets |
Learn the concepts in four minutes: what a benchmark is, what an eval adds, how evidence moves from market comparison to bank approval, and how to challenge a headline score.
Open the full glossary →A benchmark is a standardised test: the same tasks, conditions and scoring method are applied to multiple models or systems.
Use public benchmarks to shortlist models and expose obvious weaknesses.
Understand this layer →02Use-case golden setTest representative, difficult and high-risk cases drawn from the bank's work.
Understand this layer →03Controls & adversarialProbe leakage, entitlements, prompt injection, conduct, bias and unsafe actions.
Understand this layer →04Operational fitnessMeasure latency, cost, availability, fallbacks, auditability and human escalation.
Understand this layer →05Continuous evaluationRerun on model, prompt, data, tool and policy changes; monitor production drift.
Understand this layer →Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.
Primary sources linked| Compare | |||||||
|---|---|---|---|---|---|---|---|
| Public benchmark | Finance & investment, Legal & compliance, Operations & agents | Emerging | Open dataset | 2026-07-26 | Mercor Open ↗ | ||
| Internal test pack | Marketing & customer, Legal & compliance | Build internally | Internal | 2026-07-25 | Institution-owned Open ↗ | ||
| Internal test pack | Procurement & third party, Legal & compliance, Operations & agents | Build internally | Internal | 2026-07-25 | Institution-owned Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open source | 2026-07-26 | Handshake AI Research and collaborators Open ↗ | ||
| Public benchmark | Marketing & customer, Operations & agents | Established | Open dataset | 2026-07-26 | PolyAI Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents, Knowledge & RAG | Emerging | Open source | 2026-07-26 | Rogo Technologies Open ↗ | ||
| Public benchmark | Legal & compliance, Knowledge & RAG | Growing | Public methodology | 2026-07-26 | Harvey Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open dataset | 2026-07-26 | HiThink Research Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Emerging | Open source | 2026-07-26 | Longitude Labs and collaborators Open ↗ | ||
| Public benchmark | Legal & compliance, Procurement & third party | Established | Open dataset | 2026-07-25 | Stanford NLP Open ↗ | ||
| Public benchmark | Knowledge & RAG | Growing | Open source | 2026-07-25 | Meta and collaborators Open ↗ | ||
| Public benchmark | Marketing & customer, Operations & agents | Growing | Public methodology | 2026-07-25 | Salesforce AI Research Open ↗ | ||
| Public benchmark | Legal & compliance, Procurement & third party | Established | Open dataset | 2026-07-25 | The Atticus Project Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents, Knowledge & RAG | Growing | Open source | 2026-07-26 | Vals AI Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG | Established | Open dataset | 2026-07-25 | Patronus AI Open ↗ | ||
| Public benchmark | Finance & investment | Growing | Open dataset | 2026-07-26 | FinanceReasoning research team Open ↗ | ||
| Public benchmark | Finance & investment | Growing | Open source | 2026-07-25 | The FinAI Open ↗ | ||
| Public benchmark | Finance & investment | Established | Open dataset | 2026-07-25 | IBM Research and collaborators Open ↗ | ||
| Public benchmark | Legal & compliance, Operations & agents, Knowledge & RAG | Emerging | Open source | 2026-07-26 | Harvey Open ↗ | ||
| Public benchmark | Legal & compliance | Established | Open dataset | 2026-07-25 | Stanford and legal contributors Open ↗ | ||
| Public benchmark | Legal & compliance | Growing | Open dataset | 2026-07-26 | LEXam research consortium Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG, Operations & agents | Emerging | Open source | 2026-07-26 | Databricks Open ↗ | ||
| Public benchmark | Knowledge & RAG | Growing | Open dataset | 2026-07-25 | Galileo Open ↗ | ||
| Public benchmark | Finance & investment, Operations & agents | Growing | Open source | 2026-07-26 | SpreadsheetBench contributors Open ↗ | ||
| Public benchmark | Finance & investment, Knowledge & RAG | Established | Open dataset | 2026-07-25 | NExT++ Open ↗ | ||
| Public benchmark | Operations & agents, Procurement & third party | Growing | Open source | 2026-07-25 | ServiceNow Research Open ↗ | ||
| Public benchmark | Operations & agents, Marketing & customer | Growing | Open source | 2026-07-25 | Sierra Research Open ↗ |
Patronus AI
Open-book financial question answering over public-company filings with evidence-linked answers.
Test research assistants, filing analysis, credit research and finance RAG before adding institution-specific documents.
Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.
Updated with the main intelligence pipelineNew research benchmarks the serving costs of agentic memory frameworks against standard context window strategies over long conversations.
Open source ↗Research shows specialist agent decomposition outperforms monolithic LLM prompting in complex European listed real estate financial analysis.
Open source ↗Researchers applied GRPO reinforcement learning to fine-tune an LLM for financial advice, reportedly beating frontier commercial models.
Open source ↗Research proposes trajectory-adapted uncertainty quantification for LLM agents to track multi-step error propagation across tool calls.
Open source ↗Researchers introduce TradingMoE, a Mixture-of-Experts routing framework designed for LLM-based trading across changing market conditions.
Open source ↗OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.
The daily briefing flags material benchmark and assurance developments. For catalogue governance or team access, speak to OneBench directly.
Free. Daily at 06:30 UK. Unsubscribe with one click.
| 2assets |
| 4assets |
| 0assets |
| Marketing & customer | 3assets | 1assetLocal set required | 1asset | 2assets | 0assets |
|---|
| Procurement & third party | 3assets | 1assetLocal set required | 1asset | 2assets | 0assets |
|---|
| Operations & agents | 13assets | 1assetLocal set required | 1asset | 13assets | 0assets |
|---|
| Knowledge & RAG | 9assets | 0assetsLocal set required | 0assets | 5assets | 0assets |
|---|
ProForma-20Q benchmark introduces multi-period financial statement forecasting across 78 line items up to 20 quarters ahead.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion