FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition
FairTree, a new algorithm, offers subgroup fairness auditing for ML models, addressing continuous covariates better than SliceFinder/SliceLine.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
FairTree, a new algorithm, offers subgroup fairness auditing for ML models, addressing continuous covariates better than SliceFinder/SliceLine.
Research claims LLMs detect incorrectness but agree with user's false beliefs due to 'sycophancy-lying circuit' in attention heads.
Research identifies 'ungrounded reasoning' in LLMs where models fabricate answers due to lacking inferential boundary awareness, not reasoning capability.
Research proposes detoxifying large language model pre-training datasets to fundamentally reduce inherent model toxicity, rather than relying on post-training or inference-time methods.
Research identifies pervasive verbal tics (e.g., 'That's a great question!') in frontier LLMs, linked to RLHF and Constitutional AI alignment.
CulturALL introduces a new benchmark for evaluating LLM multilingual and multicultural competence on grounded, real-world tasks, beyond generic language.
Research demonstrates continual pre-training of smaller LLMs on specialized German medical data closes performance gap with larger general models.
Research analyzed 15 LLMs across 8 tasks to understand mechanisms driving LLM-guided evolutionary optimization, finding zero-shot ability correlates with final optimization.
Research evaluates prompt design and model selection on LLM accuracy predicting experience ratings from open-ended survey text.
Research claims indistinguishability metrics are insufficient for preventing data extraction from LLM APIs, formalizing a privacy game separation.
Research proposes a component-wise evaluation framework for medical Q&A LLMs, moving beyond semantic similarity to assess accuracy and health equity risks.
Research claims harmful intent is geometrically recoverable as linear directions or angular deviation in LLM residual streams across 12 models.
Research explores EVPO, an adaptive critic method for LLM post-training, aiming to balance variance reduction with noise in sparse-reward settings.
Research evaluates trade-offs between accuracy and energy consumption in text classification inference for LLMs.
Research identifies hybrid LLM architectures combining self-attention and state space models (e.g., Mamba) for long-context efficiency.
Research identifies 'tool-induced reasoning hallucinations' in LLMs using Code Interpreter, where models substitute tool outputs for coherent reasoning.
Research explores conditions where LLM-based verification improves solution quality over standalone LLM solvers, analyzing cost-benefit.
Research proposes framework to evaluate LLM representativeness beyond marginal response distributions, focusing on latent structures for cultural alignment.
Research finds prompt order (context-question-options vs. question-options-context) significantly impacts LLM performance in multiple-choice Q&A.
Research paper proposes Hybrid Document-Routed Retrieval (HDRR) to improve RAG robustness in financial documents by combining chunk-based retrieval with LLM-driven semantic file routing.
Research introduces CASS, a dataset and model for cross-architecture GPU code transpilation (CUDA to HIP, SASS to RDNA3), enabling learning-based translation.
Research proposes "Council Mode" multi-agent consensus to mitigate hallucination and bias in LLMs, particularly in Mixture-of-Experts architectures.
Research paper argues LLM watermarking adoption is hindered by misaligned incentives between providers, platforms, and users, citing competitive risk and governance.
Research evaluates multi-generation sampling for detecting jailbreaks in LLMs, testing lexical and generation inconsistency methods on various models.
Research suggests AI models evaluating other AI models (LVLM judges) may not generalize well across non-English languages.
Research identifies counterfactual unfairness in LLMs by testing response changes when speaker/addressee identities are swapped in humorous contexts.
Research proposes framework to test LLM sensitivity to subtle semantic changes in document comparison for 'needle-in-a-haystack' problems.
LegalBench-BR introduced as the first public benchmark for Brazilian legal decision classification, using 3,105 appellate proceedings.
Research tested 40+ prompt variants for LLM mathematical reasoning, finding a 'single-prompt ceiling' limiting complex problem-solving.
MORPHOGEN benchmark evaluates multilingual LLMs' handling of grammatical gender and morphological agreement in morphologically rich languages.
Research finds significant differences in how LLMs (GPT vs. Claude) handle multi-turn repair in dialogues, impacting reliability.
Research paper proposes LePREC, a classification approach for legal issue identification using LLMs on Malaysian court cases, extracted with GPT-4o.
Research introduces M²CQA, a benchmark for multilingual vision-language models (VLMs) exposing 'counterfactual hallucination' in culturally specific contexts.
Research explores small language models (SLMs) within agentic systems to overcome individual limitations and reduce compute, latency, and privacy risks.
Research identifies a new 'draft-based co-authoring jailbreak' vulnerability in LLMs, where incomplete drafts can compel harmful content generation.
IndiaFinBench is a new public benchmark evaluating LLM performance on Indian financial regulatory text, addressing a gap in non-Western financial NLP.
Research shows LLM personalization via sociodemographic cues can amplify biases depending on prompt phrasing and contextual cues.
Research identifies implicit local and global biases in multilingual LLMs when answering locale-ambiguous questions, creating LocQA benchmark.
RepIt, a new framework, selectively suppresses language model refusal on targeted concepts, improving upon existing steering methods.
Research paper proposes a neurosymbolic architecture (Foundation AgenticOS) for enterprise agents to address LLM hallucination and regulatory compliance via ontologies.
Research surveys dynamic model routing and cascading strategies for LLM inference to optimize performance and cost by selecting models based on query complexity.
Research paper audits information leakage in privacy-preserving in-context learning (ICL) methods, identifying potential vulnerabilities.
Research highlights that current LLM evaluation, focused on accuracy, overlooks critical enterprise factors: energy, latency, hardware utilization, and cost control.
Research proposes a novel method, 'Soft-Hybrid Alphabet Estimation,' for quantifying LLM uncertainty and unmasking hallucinations with limited query samples.
Research proposes unsupervised weight monitoring for fine-tuned LLMs to detect out-of-distribution threats like backdoors without training data access.
Researchers introduced Visual-TableQA, a large-scale, open-domain multimodal dataset and benchmark for reasoning over rendered table images.
Research finds vision-language models struggle with negation in multiple languages, exhibiting affirmation bias beyond English.
STAR-Teaming introduces a black-box, multi-agent system for automated red teaming of LLMs to generate jailbreak prompts effectively.
RARE proposes a new RAG evaluation framework for corpora with high document similarity, addressing a gap in existing benchmarks.
Research demonstrates LLM answers vary significantly based on retrieved document order in RAG, even when gold document is present.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion