1BOneBench

Search OneBench

Search signals, briefings, benchmarks and glossary terms.

Catalogue reviewed 26 Jul 202632 evaluation assets across 10 bank domains2 known public-coverage gapsLast ingest 31 Jul, 20:36 UK
The OneBench evaluation position

A leaderboard is a screening tool, not a model approval.

What a bank needs

Combine external benchmarks with a versioned use-case golden set, adversarial control tests, human baselines, operational fitness and continuous evaluation against the exact system being deployed.

Evaluation coverage matrix

Start here: read across a bank function to see where reusable evaluation evidence exists—and where public benchmarks stop. Every cell opens the catalogue filtered to that function and evaluation layer.

Amber always requires institution-owned cases
Evaluation asset coverage by bank function and evaluation-stack layer
Bank functionExternal screenPublic benchmarks and standards used to shortlist systems.Golden setRepresentative institution-owned cases for the intended use.ControlsSafety, conduct, security and policy-boundary testing.OperationalWorkflow, tool-use, cost, latency and human-escalation evidence.ContinuousRepeatable regression and production-change evaluation.
Finance & investment13assets0assetsLocal set required0assets8assets0assets
Legal & compliance7assets2assetsLocal set required2assets4assets0assets
Marketing & customer3assets1assetLocal set required1asset2assets0assets
Procurement & third party3assets1assetLocal set required1asset2assets0assets
Operations & agents13assets1assetLocal set required1asset13assets0assets
Knowledge & RAG9assets0assetsLocal set required0assets5assets0assets
Cybersecurity1asset0assetsLocal set required1asset0assets0assets
Software engineering1asset0assetsLocal set required0assets1asset0assets
Safety & governance4assets1assetLocal set required5assets2assets1asset
Model & infrastructure2assets0assetsLocal set required2assets2assets1asset
No mapped asset Some coverage Deeper coverage Local golden set required
Cybersecurity: thin by design

This catalogue covers AI-specific security evaluation. General cyber control and penetration-testing suites remain outside its scope.

Software engineering: thin by design

Coverage is limited to evaluation of AI coding systems, not conventional software-development lifecycle test suites.

Benchmarks and evals 101

Learn the concepts in four minutes: what a benchmark is, what an eval adds, how evidence moves from market comparison to bank approval, and how to challenge a headline score.

Open the full glossary →
The comparison layer

What is a benchmark?

A benchmark is a standardised test: the same tasks, conditions and scoring method are applied to multiple models or systems.

  • Useful for comparing options on a defined capability.
  • Strongest when the data, harness and scoring are inspectable.
  • Does not prove performance on your users, data or controls.
Open the glossary definition →
A bank-grade evaluation stack
Benchmark and framework catalogue

Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.

Primary sources linked
9 matching evaluation assets0 selected for comparison
Compare
Public benchmarkFinance & investment, Operations & agents, Knowledge & RAGEmergingOpen source2026-07-26Rogo Technologies
Open ↗
Public benchmarkLegal & compliance, Knowledge & RAGGrowingPublic methodology2026-07-26Harvey
Open ↗
Public benchmarkKnowledge & RAGGrowingOpen source2026-07-25Meta and collaborators
Open ↗
Public benchmarkFinance & investment, Operations & agents, Knowledge & RAGGrowingOpen source2026-07-26Vals AI
Open ↗
Public benchmarkFinance & investment, Knowledge & RAGEstablishedOpen dataset2026-07-25Patronus AI
Open ↗
Public benchmarkLegal & compliance, Operations & agents, Knowledge & RAGEmergingOpen source2026-07-26Harvey
Open ↗
Public benchmarkFinance & investment, Knowledge & RAG, Operations & agentsEmergingOpen source2026-07-26Databricks
Open ↗
Public benchmarkKnowledge & RAGGrowingOpen dataset2026-07-25Galileo
Open ↗
Public benchmarkFinance & investment, Knowledge & RAGEstablishedOpen dataset2026-07-25NExT++
Open ↗
Public benchmark

FinanceBench

Patronus AI

Established

Open-book financial question answering over public-company filings with evidence-linked answers.

Finance & investmentKnowledge & RAG
Access
Open dataset
Last reviewed
2026-07-25
Measures
4 dimensions

What it measures

  • Financial document retrieval
  • Numerical reasoning
  • Answer correctness
  • Evidence grounding

How a G-SIB should use it

Test research assistants, filing analysis, credit research and finance RAG before adding institution-specific documents.

Useful metrics

Exact / judged answer accuracyEvidence retrievalCitation correctness
Primary sourceOpen FinanceBench
Evaluation and benchmark watch

Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.

Updated with the main intelligence pipeline
arXiv cs.LG — Machine Learning

Dynamically Scaled Activation Steering

Researchers introduced Dynamically Scaled Activation Steering (DSAS), a method-agnostic framework for guiding generative model behavior more efficiently.

Open source ↗
arXiv cs.LG — Machine Learning

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning

Research shows averaging evaluation metrics masks significant capability trade-offs in LLM pruning, particularly in code retention.

Open source ↗
arXiv cs.LG — Machine Learning

Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards

New research shows previous claims of large language models learning effectively from 100% noisy data using RLVR are invalid due to data contamination.

Open source ↗
arXiv cs.LG — Machine Learning

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Researchers introduced Echoverse, a framework for generating deep, evolving synthetic environments to train computer-use agents at scale.

Open source ↗
arXiv cs.LG — Machine Learning

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Research investigates deep generative models' ability to reproduce complex, non-stationary spatial and spatio-temporal data distributions, a key challenge for real-world application.

Open source ↗
arXiv cs.LG — Machine Learning

Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection

Research introduces a 'three-number reporting' method for evaluating classifiers by matching evaluation conditions to operational class imbalance for Sentinel-1 wave detection.

Open source ↗
arXiv cs.LG — Machine Learning

Uncertainty quantification for trustworthy deep learning: Methods and measures

A research survey reviews methods for Uncertainty Quantification (UQ) in deep learning, focusing on ensemble and approximate Bayesian approaches.

Open source ↗
arXiv cs.LG — Machine Learning

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

Research finds Diffusion Language Models (DLMs) are less robust to input noise and adversarial attacks than autoregressive (AR) models like Llama 3.

Open source ↗
arXiv cs.LG — Machine Learning

Error Analysis of Neural-Network-Based Engression

Research presents a theoretical error analysis for Engression, a neural-network-based conditional distribution learning method using the energy score.

Open source ↗
arXiv cs.LG — Machine Learning

STEREODISCO: Discovering Stereotypicality in LLMs

STEREODISCO framework measures and discovers previously unexamined stereotypical associations and biases in LLMs using semantic differential method.

Open source ↗
arXiv cs.LG — Machine Learning

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

Research introduces Conformal Cascade, a multi-tier LLM inference method offering distribution-free accuracy guarantees, addressing miscalibrated confidence scores.

Open source ↗
arXiv cs.LG — Machine Learning

Distributions In, Distributions Out: The Case for Soft-Label Training

Research proposes 'soft-label training' using a distribution of annotator judgments instead of single majority labels for ambiguous tasks.

Open source ↗
Methodology

OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.

Keep pace with finance AI evaluation

The daily briefing flags material benchmark and assurance developments. For catalogue governance or team access, speak to OneBench directly.