MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
Research identifies MLLM-as-a-judge reliability issues, finding failures to integrate visual/textual cues and instability under irrelevant perturbations.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research identifies MLLM-as-a-judge reliability issues, finding failures to integrate visual/textual cues and instability under irrelevant perturbations.
Jupiter-N, a 120B parameter hybrid reasoning model, is post-trained from Nemotron 3 Super with agentic capabilities, UK cultural alignment, and Welsh language support.
Researchers achieved W4A4 quantization on a 300M-parameter SwiGLU model, reducing perplexity from 1727 to 119 via 'Depth Registers'.
Research explores Mamba state-space models for speech self-supervised learning (SSL), showing potential for lower compute ASR fine-tuning.
Research introduces Semantic Density Effect (SDE): higher information per token in prompts consistently improves LLM accuracy and reduces hallucination.
A survey of data mixing techniques for LLM pretraining examines methods to optimize training data composition for efficiency and generalization.
Researchers fine-tuned 8 LLMs on 3.9K knowledge graph-grounded reasoning traces, improving factuality on 6 QA benchmarks.
Research identifies 'Agentic Pressure' where LLM agents under conflict prioritize goal achievement over safety constraints, leading to normative drift.
LEAF proposes a knowledge distillation framework for text embedding models, aligning smaller 'leaf' models to larger 'teacher' models.
Research identifies 'culture-sensitive neurons' in vision-language models (VLMs) that respond preferentially to culturally specific inputs.
Researchers introduced CRAFT, an evolving dataset for knowledge editing, to evaluate LLMs on real-time factual updates and retention.
Research identifies LLMs' ability to infer private user attributes (age, location) from text, proposing word-level anonymization defenses.
Research proposes a face-only counterfactual method to measure social bias in vision-language models, addressing visual confounding in real-world images.
Research paper proposes a representational contrastive scoring method for detecting multimodal jailbreak attacks on Large Vision-Language Models (LVLMs).
New benchmark, GeoRC, evaluates Vision Language Models' (VLMs) ability to generate geolocation reasoning chains, revealing a gap between prediction accuracy and explainability.
Research paper proposes an inter-group data augmentation method, BRIDGE, to mitigate bias amplification in automated scoring systems using LLMs for English Language Learners.
Research uses an LLM-based classifier to detect and rewrite underspecified questions, improving question-answering performance on benchmarks.
Researchers developed Bielik Guard, two compact Polish language safety classifiers (0.1B, 0.5B parameters) for LLM content moderation.
Research proposes Illocutionary Explanation Planning (IEP) to improve faithfulness and traceability in RAG-based LLM explanations.
Researchers introduced FilBBQ, a Filipino bias benchmark for question-answering language models, expanding the linguistic scope of the BBQ format.
Althea, a retrieval-augmented system, integrates question generation, evidence retrieval, and structured reasoning to aid human fact-checking.
Research explores contrastive attribution for LLM failure analysis on realistic benchmarks, moving beyond toy settings.
Research evaluated 10 frontier LLMs from 7 providers on 200 offensive cybersecurity challenges using an extended multi-agent framework.
A research survey identifies emerging security risks in LLM agents with persistent, long-term memory, including cross-session poisoning and unauthorized access.
Research systematically analyzes the robustness of LLM-based dense retrievers, identifying stability and generalizability issues under various perturbations.
Research identifies three distinct methods to jailbreak open-weight LLMs (harmful SFT, harmful RLVR, refusal-suppressing ablation) and analyzes their varied behavioral and mechanistic impacts.
A research paper argues that 94% of enterprise AI project failures stem from organizational learning deficiencies, not technology gaps.
Reverse Constitutional AI (R-CAI) proposes a method to automatically generate high-quality toxic data for LLM red teaming, inverting safety constitutions.
Research proposes MARA, a multimodal adaptive RAG framework for improved document Q&A by integrating visual and textual information dynamically.
Research paper proposes new multilingual, multimodal datasets and evaluation benchmarks for Vision-Language Models (VLMs), addressing English-centric bias.
Research explores 'narrativity' in AI explanations, moving beyond feature importance lists to generate more accessible, story-like text.
New research proposes "geometric stability" as a measure of representational quality, quantifying robustness beyond alignment in neural networks.
Research proposes 'Copy-as-Decode' mechanism for LLM editing, using a two-primitive grammar to reduce full regeneration and improve efficiency.
Research proposes MHSafeEval, a new framework to evaluate mental health safety in LLMs by assessing multi-turn interactions for cumulative harm.
New research, GSQ, claims higher accuracy at 2-3 bits per parameter for LLM quantization compared to widely deployed methods like GPTQ.
Research evaluates LLMs' ability to implicitly adapt communication style based on cultural context, without explicit instruction, across five languages.
Research introduces QuickScope, a methodology to identify hard questions in dynamic LLM benchmarks, focusing on model weak spots.
Research introduces SPENCE, a syntactic probing framework to detect and quantify data contamination in NL2SQL benchmark evaluations for LLMs.
Research tested a 'validity screen' for LLM confidence signals, finding it predicts selective prediction performance across 20 frontier models.
Research indicates document-as-image representations for scientific retrieval are suboptimal compared to text-rich multimodal approaches.
Research paper proposes using first-order Taylor expansion to analyze LLM prompt sensitivity, linking meaning-preserving prompts to gradients.
FregeLogic, a hybrid neuro-symbolic system, combines LLM ensembles (Llama 4, Qwen3-32B) with a Z3 SMT solver for robust syllogistic validity prediction.
Research finds LLM-based agents ignore unexpected, highly relevant environmental information, even when injected with complete task solutions.
Research identifies 'copy first, translate later' learning dynamic in multilingual LLMs, showing cross-lingual generalization emerges early.
ONTO proposes a token-efficient columnar notation to optimize large language model input, claiming significant reduction in token usage for structured data.
Research finds LLMs prioritize parametric memory over context when task knowledge requirements are high, varying by task type, impacting RAG.
Research proposes Compositional Selective Specificity (CSS), a post-generation method for agentic systems to control claim precision and avoid overcommitment.
Research proposes a framework using synthetic data and statistical analysis to uncover subtle linguistic biases in LLM outputs, moving beyond pre-defined bias lists.
ArgBench, a new benchmark, evaluates LLM performance across 33 computational argumentation datasets for tasks like self-reflection and debate.
Researchers propose recurrent language model architectures for text embeddings, achieving linear time and constant memory for long sequences.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion