CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
Research introduces CompliBench, a benchmark for evaluating LLM judges' ability to detect compliance violations in dialogue systems.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces CompliBench, a benchmark for evaluating LLM judges' ability to detect compliance violations in dialogue systems.
Research identifies visual token dominance as the core bottleneck in large Vision-Language Model (LVLM) inference efficiency, proposing a taxonomy of techniques.
New arXiv paper proposes benchmarks for Large Vision-Language Models (LVLMs) to test deflection and hallucination with conflicting visual and textual evidence.
AlphaEval proposes a new framework for evaluating AI agents in production environments, accounting for heterogeneous, multi-modal inputs and implicit constraints.
Research proposes 'reasoning calibration' to improve LLM factuality in long-form generation by enabling models to estimate reliability of claims.
Research finds LLMs exhibit the 'Identifiable Victim Effect,' prioritizing narratively described individuals over statistically larger groups in resource allocation.
ReasonXL paper claims LLMs can be fine-tuned to reason in non-English languages without performance loss, addressing English-centric reasoning.
Researchers propose ToxiTrace, a BERT-style model method using LLM guidance for explainable Chinese toxic content detection with fine-grained toxic span identification.
Research proposes a novel method, GRADE, using gradient subspace dynamics to probe LLM internal knowledge gaps, aiming for better confidence detection.
Research introduces Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework to improve LLM precision and reproducibility in text categorization.
Research proposes multi-agent, multi-format approach for LLMs to understand complex spreadsheets, addressing layout cues and scale limits.
Research explores using LLMs to evaluate data privacy and AI safety in contexts with imperfect information, moving beyond complete context assumptions.
Research explores if LLMs possess 'privileged knowledge' about their own answer correctness from internal states, beyond external observation.
Research paper proposes "SeedPrints" method to identify the random seed used to train a Large Language Model for provenance and attribution.
New arXiv research questions if VLMs genuinely understand candlestick charts for stock forecasting, citing inadequate benchmarks.
Researchers demonstrated a clean-label backdoor attack on Graph Neural Networks (GNNs), manipulating predictions without altering training node labels.
Research paper proposes Parcae, a new training recipe for stable, looped language models that scales quality via recurrent computation within fixed parameters.
Research explores Monte Carlo Stochastic Depth (MCSD) to enhance uncertainty quantification (UQ) in deep learning, building on MC Dropout methods.
Research proposes provably replicable reinforcement learning algorithms with linear function approximation to address experimental variability.
Research proposes Calibration-Aware Policy Optimization (CAPO) to improve LLM reasoning calibration, addressing overconfidence from GRPO-style algorithms.
Nemotron 3 Super, a 120B parameter hybrid Mamba-Attention Mixture-of-Experts model, introduces NVFP4 pre-training and LatentMoE architecture.
Research introduces INTARG, a new method for generating real-time adversarial attacks on time-series regression models, impacting forecasting systems.
Research explores feature disentanglement to mitigate 'shortcut learning' in deep learning models, improving generalization by reducing reliance on spurious correlations.
New research introduces CodeRQ-Bench, a benchmark for evaluating LLM reasoning quality across various coding tasks beyond just code generation.
Research proposes design-time verification for AI models to ensure numerical stability, computational correctness, and domain consistency before training.
Research identifies 'semantic fixation' in VLMs: models default to familiar interpretations despite explicit prompt instructions, impacting rule-mapping. New VLM-Fix benchmark introduced.
Research identifies 'policy-invisible violations' in LLM agents, where valid actions violate hidden organizational policies due to missing context.
Research introduces SpecBound, a speculative decoding method for LLMs using self-drafting with layer-wise confidence calibration to improve inference speed.
Research proposes LLM-guided semantic bootstrapping to transfer LLM knowledge into interpretable Tsetlin Machines for text classification.
FaCT (Faithful Concept Traces) proposes a new concept-based interpretability method for neural networks, aiming for improved faithfulness and fewer assumptions.
RankOOD proposes a new Out-of-Distribution (OOD) detection method using Placket-Luce loss for training, leveraging ranking patterns in ID class predictions.
Research demonstrates backdoors can be embedded into AI agent fine-tuning data pipelines, leading to malicious behavior upon trigger.
Researchers failed to reliably distill behavioral dispositions (self-verification, uncertainty) into small language models (0.6B-2.3B parameters).
Researchers introduced INDOTABVQA, a benchmark for cross-lingual Table Visual Question Answering (VQA) in Bahasa Indonesia documents.
Research paper proposes "face density" as a quantifiable metric for data complexity in machine learning, beyond simple instance count.
New research proposes a bootstrap method for uncertainty quantification in Convolutional Neural Networks (CNNs), addressing a gap in theoretical consistency.
Research finds stronger reasoning LLMs can reduce fidelity in behavioral simulations when the goal is to sample boundedly rational behavior, not solve problems.
Research shows multi-token prediction (MTP) consistently outperforms next-token prediction (NTP) for planning tasks in Transformers.
Research identifies key conditions for successful on-policy distillation of LLMs, focusing on student-teacher thinking pattern compatibility.
Research benchmarks LLM-enhanced log anomaly detection against traditional methods for system diagnostics, demonstrating potential for operational reliability.
New research introduces "Socrates Loss," a single-loss function to improve confidence calibration and classification in deep neural networks, addressing a key trade-off.
Research finds Transformer and LLM models can infer applicant gender from academic recommendation letters even with explicit identifiers removed, due to implicit language patterns.
New research proposes Shortcut Guardrail, a deployment-time framework to mitigate token-level shortcut learning in language models without retraining.
Research claims fundamental limits in verifying AI model calibration, stating that error rates below a statistical noise floor are unmeasurable.
Researchers propose BID-LoRA, a parameter-efficient framework combining continual learning (CL) and machine unlearning (MU) capabilities.
Research analyzes the effect of various noise types in fine-tuning datasets on LLM performance and proposes methods to mitigate degradation.
Researchers propose Outlier Separation in Channel (OSC) for W4A4 quantization, improving 4-bit LLM inference accuracy by addressing activation outliers.
Research identifies large language models (LLMs) exhibit safety vulnerabilities in low-resource languages due to biased safety alignment.
GF-Score proposes a framework to evaluate class-conditional adversarial robustness for neural networks, decomposing certified scores into per-class profiles.
Notion cofounder and Head of AI discuss their journey shipping AI agents for knowledge work, detailing multiple rebuilds and tool integrations.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion