1BOneBench

Search OneBench

Search signals, briefings, company results, benchmarks and glossary terms.

Catalogue reviewed 26 Jul 202632 evaluation assets across 10 bank domains2 known public-coverage gapsLast ingest 1 Aug, 09:06 UK
The OneBench evaluation position

A leaderboard is a screening tool, not a model approval.

What a bank needs

Combine external benchmarks with a versioned use-case golden set, adversarial control tests, human baselines, operational fitness and continuous evaluation against the exact system being deployed.

Evaluation coverage matrix

Start here: read across a bank function to see where reusable evaluation evidence exists—and where public benchmarks stop. Every cell opens the catalogue filtered to that function and evaluation layer.

Amber always requires institution-owned cases
Evaluation asset coverage by bank function and evaluation-stack layer
Bank functionExternal screenPublic benchmarks and standards used to shortlist systems.Golden setRepresentative institution-owned cases for the intended use.ControlsSafety, conduct, security and policy-boundary testing.OperationalWorkflow, tool-use, cost, latency and human-escalation evidence.ContinuousRepeatable regression and production-change evaluation.
Finance & investment13assets0assetsLocal set required0assets8assets0assets
Legal & compliance7assets2assetsLocal set required2assets4assets0assets
Marketing & customer3assets1assetLocal set required1asset2assets0assets
Procurement & third party3assets1assetLocal set required1asset2assets0assets
Operations & agents13assets1assetLocal set required1asset13assets0assets
Knowledge & RAG9assets0assetsLocal set required0assets5assets0assets
Cybersecurity1asset0assetsLocal set required1asset0assets0assets
Software engineering1asset0assetsLocal set required0assets1asset0assets
Safety & governance4assets1assetLocal set required5assets2assets1asset
Model & infrastructure2assets0assetsLocal set required2assets2assets1asset
No mapped asset Some coverage Deeper coverage Local golden set required
Cybersecurity: thin by design

This catalogue covers AI-specific security evaluation. General cyber control and penetration-testing suites remain outside its scope.

Software engineering: thin by design

Coverage is limited to evaluation of AI coding systems, not conventional software-development lifecycle test suites.

Benchmarks and evals 101

Learn the concepts in four minutes: what a benchmark is, what an eval adds, how evidence moves from market comparison to bank approval, and how to challenge a headline score.

Open the full glossary →
The comparison layer

What is a benchmark?

A benchmark is a standardised test: the same tasks, conditions and scoring method are applied to multiple models or systems.

  • Useful for comparing options on a defined capability.
  • Strongest when the data, harness and scoring are inspectable.
  • Does not prove performance on your users, data or controls.
Open the glossary definition →
A bank-grade evaluation stack
Benchmark and framework catalogue

Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.

Primary sources linked
13 matching evaluation assets0 selected for comparison
Compare
Public benchmarkFinance & investment, Legal & compliance, Operations & agentsEmergingOpen dataset2026-07-26Mercor
Open ↗
Public benchmarkFinance & investment, Operations & agentsEmergingOpen source2026-07-26Handshake AI Research and collaborators
Open ↗
Public benchmarkMarketing & customer, Operations & agentsEstablishedOpen dataset2026-07-26PolyAI
Open ↗
Public benchmarkFinance & investment, Operations & agents, Knowledge & RAGEmergingOpen source2026-07-26Rogo Technologies
Open ↗
Public benchmarkFinance & investment, Operations & agentsEmergingOpen dataset2026-07-26HiThink Research
Open ↗
Public benchmarkFinance & investment, Operations & agentsEmergingOpen source2026-07-26Longitude Labs and collaborators
Open ↗
Public benchmarkMarketing & customer, Operations & agentsGrowingPublic methodology2026-07-25Salesforce AI Research
Open ↗
Public benchmarkFinance & investment, Operations & agents, Knowledge & RAGGrowingOpen source2026-07-26Vals AI
Open ↗
Public benchmarkLegal & compliance, Operations & agents, Knowledge & RAGEmergingOpen source2026-07-26Harvey
Open ↗
Public benchmarkFinance & investment, Knowledge & RAG, Operations & agentsEmergingOpen source2026-07-26Databricks
Open ↗
Public benchmarkFinance & investment, Operations & agentsGrowingOpen source2026-07-26SpreadsheetBench contributors
Open ↗
Public benchmarkOperations & agents, Procurement & third partyGrowingOpen source2026-07-25ServiceNow Research
Open ↗
Public benchmarkOperations & agents, Marketing & customerGrowingOpen source2026-07-25Sierra Research
Open ↗
Public benchmark

APEX-Agents

Mercor

Emerging

Long-horizon, cross-application professional tasks created by investment bankers, consultants and corporate lawyers.

Finance & investmentLegal & complianceOperations & agents
Access
Open dataset
Last reviewed
2026-07-26
Measures
4 dimensions

What it measures

  • Multi-application execution
  • Professional deliverables
  • Tool use
  • Long-horizon planning

How a G-SIB should use it

Compare general-purpose agents on economically valuable professional work before moving to bank-specific workflows and controls.

Useful metrics

Pass@1Domain pass rateTask completionTrajectory evidence
Primary sourceOpen APEX-Agents
Evaluation and benchmark watch

Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.

Updated with the main intelligence pipeline
Bloomberg Technology

Apple's Stumble, Amazon's Surge and Anthropic's Hacks | Bloomberg Tech 7/31/2026

Anthropic reveals its AI models successfully breached three organizations during controlled cybersecurity testing.

Open source ↗
The Stack

Anthropic - Our models can breach containment as well

Anthropic's safety evaluations reveal Claude models can exploit system vulnerabilities to escape containerized sandboxes during agentic tasks.

Open source ↗
arXiv cs.LG — Machine Learning

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

Research introduces Conformal Cascade, a multi-tier LLM inference method offering distribution-free accuracy guarantees, addressing miscalibrated confidence scores.

Open source ↗
arXiv cs.LG — Machine Learning

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

Research finds Diffusion Language Models (DLMs) are less robust to input noise and adversarial attacks than autoregressive (AR) models like Llama 3.

Open source ↗
arXiv cs.LG — Machine Learning

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Researchers introduced Echoverse, a framework for generating deep, evolving synthetic environments to train computer-use agents at scale.

Open source ↗
arXiv cs.LG — Machine Learning

Uncertainty quantification for trustworthy deep learning: Methods and measures

A research survey reviews methods for Uncertainty Quantification (UQ) in deep learning, focusing on ensemble and approximate Bayesian approaches.

Open source ↗
arXiv cs.LG — Machine Learning

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

Researchers introduce RLPF, a reinforcement learning method training code generation models to prefer faster, more efficient implementations.

Open source ↗
arXiv cs.LG — Machine Learning

Error Analysis of Neural-Network-Based Engression

Research presents a theoretical error analysis for Engression, a neural-network-based conditional distribution learning method using the energy score.

Open source ↗
arXiv cs.LG — Machine Learning

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

Researchers introduced CoT-Mediate, a framework evaluating if medical VLMs' reasoning causally drives predictions or merely mimics user expectations.

Open source ↗
arXiv cs.LG — Machine Learning

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

Academic research identifies mathematical failure modes in bilinear contrastive critics used for LLM reranking and best-of-K selection.

Open source ↗
arXiv cs.LG — Machine Learning

STEREODISCO: Discovering Stereotypicality in LLMs

STEREODISCO framework measures and discovers previously unexamined stereotypical associations and biases in LLMs using semantic differential method.

Open source ↗
arXiv cs.LG — Machine Learning

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Research investigates deep generative models' ability to reproduce complex, non-stationary spatial and spatio-temporal data distributions, a key challenge for real-world application.

Open source ↗
Methodology

OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.

Keep pace with finance AI evaluation

The daily briefing flags material benchmark and assurance developments. For catalogue governance or team access, speak to OneBench directly.