1BOneBench

Search OneBench

Search signals, briefings, benchmarks and glossary terms.

Friday, 31 July 2026180 qualifying developments across 2 sourcesLast ingest 31 Jul, 20:36 UK
Clear all
Investigate12 results
Research radar

Latest research worth inspecting

Preprints and research are separated from the executive feed. Publication here is not validation; open the paper and inspect its evidence.

Showing 12 of 180 · latest first
ResearcharXiv cs.LG — Machine Learning

STEREODISCO: Discovering Stereotypicality in LLMs

Open source ↗
Executive summary

STEREODISCO framework measures and discovers previously unexamined stereotypical associations and biases in LLMs using semantic differential method.

InvestigateNext 12 months
model biasresponsible aimodel evaluationsafety alignmentllm security
ResearcharXiv cs.LG — Machine Learning

Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

Open source ↗
Executive summary

Research proposes Doubly Robust Functional Representation Learning for longitudinal causal inference, handling irregular time-series data in studies.

InvestigateNext 12 months
causal inferencelongitudinal studiesirregular time seriesrepresentation learningmachine learning research
Show interpretive assessment

So whatImproved causal inference with irregular time-series data could enhance credit risk and health economics models.

Do whatBrief your quantitative research team on this method's potential for robust causality in high-stakes time-series analysis.

ResearcharXiv cs.LG — Machine Learning

Uncertainty quantification for trustworthy deep learning: Methods and measures

Open source ↗
Executive summary

A research survey reviews methods for Uncertainty Quantification (UQ) in deep learning, focusing on ensemble and approximate Bayesian approaches.

InvestigateNext 12 months
uncertainty quantificationmodel riskexplainabilitysafety alignmentresponsible ai
Show interpretive assessment

So whatRobust uncertainty quantification is critical for deploying deep learning models in safety-critical financial domains.

Do whatAdd this survey to the model risk research backlog for review by your validation and responsible AI teams.

ResearcharXiv cs.LG — Machine Learning

Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models

Open source ↗
Executive summary

Research finds Diffusion Language Models (DLMs) are less robust to input noise and adversarial attacks than autoregressive (AR) models like Llama 3.

InvestigateNext 12 months
model evaluationllm securitymodel robustnessresearch trends
Show interpretive assessment

So whatDiffusion Language Models show lower robustness than AR models, complicating their use in financial services.

Do whatAdd Diffusion Language Models to your next model evaluation framework discussion.

ResearcharXiv cs.CL — Computation and Language

Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore

Open source ↗
Executive summary

Researchers introduced LogicScore, a new evaluation method for Retrieval Augmented Generation (RAG) systems focused on global logical integrity.

InvestigateNext 12 months
ragmodel evaluationexplainabilityresponsible ai
Show interpretive assessment

So whatRAG evaluations shifting beyond factual recall to logical coherence directly impacts model risk and responsible AI frameworks.

Do whatAdd LogicScore methodology to the model validation team's emerging techniques watchlist for advanced RAG evaluation.

ResearcharXiv cs.CL — Computation and Language

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

Open source ↗
Executive summary

KletterMix introduces a high-quality, reusable German pretraining corpus for language models, addressing the scarcity of well-curated German linguistic resources.

InvestigateNext 12 months
pretraining datagerman language modelscorpus developmentmodel qualityresource curation
Show interpretive assessment

So whatHigh-quality German pretraining data enables competitive local-language model performance, influencing European market strategy.

Do whatAdd KletterMix to your data strategy team's resource evaluation for German-language models.

ResearcharXiv cs.CL — Computation and Language

Ranked by Position: Order Sensitivity as an Exploitable Attack Surface in LLM Listwise Recommenders

Open source ↗
Executive summary

Research reveals LLM-based recommender systems are vulnerable to position bias, enabling attackers to promote items by reordering candidates.

InvestigateNext 12 months
llm securitymodel riskrecommendation systemssafety alignmentmodel evaluation
Show interpretive assessment

So whatPositional bias in LLM rerankers creates an attack surface for manipulating recommendation outputs in financial product surfacing.

Do whatAdd `promo@k` metrics to your model validation framework for LLM-based recommender systems.

What this board does—and does not—say

The default view excludes research papers, removes low-confidence items, sorts by publication date and caps each publisher at four displayed items. Research has its own view, capped at six papers per research feed. Every headline opens the underlying source.

One development is evidence, not momentum. The board does not label a topic “rising” from a single article, and the narrative implications are explicitly marked as interpretive assessments. Use repeated, independent sources over time before treating a topic as a trend.

Start with the decision-ready evidence

Receive eight source-linked developments at 06:30 UK, with factual summary kept separate from interpretive assessment.