ASSERT: A Measurement Pipeline for GenAI Audits
Researchers introduced ASSERT, a measurement pipeline designed to isolate system changes from audit methodology choices in GenAI evaluations.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduced ASSERT, a measurement pipeline designed to isolate system changes from audit methodology choices in GenAI evaluations.
Researchers evaluated claim-level hallucination rates across eight legal RAG systems using English GDPR and French civil law corpora.
ArXiv research demonstrates that LLM high-confidence errors exhibit stable miscalibration that persists under small perturbations.
Research proposes structural abstention for LLM text-to-SQL models to prevent silent hallucinations in enterprise data dashboards.
Researchers introduced AnchorBench to evaluate how numerical anchors and initial reference values trigger cognitive biases in LLMs.
Researchers introduce adaptive stopping mechanisms for multi-turn LLM reasoning and ReAct agents to reduce redundant inference calls.
New research shows benchmark optimization on SWE-bench and LiveCodeBench fails to reflect real-world enterprise coding capabilities.
Researchers propose a robust XGBoost regression variant to mitigate sensitivity to vertical outliers and leverage points in tabular datasets.
An arXiv paper details eight years of language model evolution, tracking rapid capability growth in coding agents alongside steep cost declines.
Researchers introduce MINT, a zero-shot foundation model architecture designed for multi-task predictive reasoning on transaction sequence data.
Researchers introduced fairness constraints directly into Tabular Foundation Model (TFM) pre-training to enable fair predictions.
Research demonstrates no universal inference-time signal can predict individual sample-level LLM regressions during version updates.
Researchers propose applying ACID-compliant transactional frameworks to LLM agents to ensure reliable, safe, and durable multi-step execution.
Researchers introduced Regime-Conditional Verification to adapt safety classifiers to policy shifts and traffic drift without retraining.
Researchers introduced a graph-based RL framework to detect, assess, and recover from runtime behavioral drift in autonomous LLM agents.
Researchers introduced Principle-Bench to evaluate LLM-as-judge systems assessing compliance with principle-based financial regulations.
Researchers introduced ATLAS, a framework using automata learning to extract interpretable behavioral strategies from complex LLM agents.
Researchers unified leading membership inference attacks (LiRA, RMIA, BASE) into a single exponential-family framework called BaVarIA.
Researchers propose a gradient-guided token suppression defense against visual prompt injection attacks on multimodal models.
Academic paper formalizes the tradeoff between training window length and model complexity in non-stationary financial return prediction.
A European bank tested a hybrid AI-econometric prototype combining sentiment analysis and topic modeling for interest rate forecasting in ALM.
Researchers introduced HalluTruthQA-4K, an expanded fine-grained evaluation corpus and dataset for Arabic hallucination detection.
A research paper demonstrates that classification models lose probability calibration when encountering unseen fine-grained subtypes.
Researchers propose a validation framework using LLM agents conditioned on behavioral profiles to simulate and vet A/B test outcomes.
Researchers introduced RepBench, a standardized benchmark dataset mapping internal LLM representations to 182 distinct capability taxonomies.
Kalypso introduces relational LLM serving, an abstraction making LLM execution aware of query plans to improve performance for semantic operations on unstructured data.
Research dissociates LLM sycophancy into factual and opinion subtypes, analyzing internal representations to understand its varied manifestations.
Research proposes a category theory approach using Kan extensions to define and compute structural invariants for transfer learning between tasks.
Research proposes a reward modeling approach to debias Text-to-Image (T2I) evaluation by incorporating implicit cultural alignment.
Microsoft delayed an Exchange update due to a backlog of bugs introduced by machine-generated code tools.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion