Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Researchers propose Contrastive Policy Optimization (CPO) for reinforcement learning with verifiable rewards, using token-level disagreement for correctness.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose Contrastive Policy Optimization (CPO) for reinforcement learning with verifiable rewards, using token-level disagreement for correctness.
Research proposes a new post-training method for multimodal document Q&A to improve visual grounding without high inference costs or large datasets.
A study found single-pass frontier models outperform multi-agent debate tools for providing useful feedback on economics meta-analyses.
Researchers introduced 'The Energy Society,' a simulation where LLM agents' survival depends on energy tied to inference cost and job completion.
Research shows finetuning LLMs on narrow, factually-defensible data can cause broad, unintended ideological shifts in unrelated domains.
Research explores instruction tuning and model merging to adapt reasoning language models for domains lacking reliable output verification.
Research introduces "delta signal" for on-policy distillation, an alternative post-training method in reinforcement learning for token-level supervision.
MedFailBench, a new clinician-built open-source benchmark, evaluates medical AI safety by categorizing errors into specific failure types and severity levels.
Research benchmarks six MLLMs on a scientific visualization literacy test, finding current evaluations are chart-centric and lack SciVis understanding evidence.
Research finds static retrieval evaluation methods fail to predict utility for multi-step LLM agents, which use documents to inform subsequent actions.
Research demonstrates that pretraining data can be poisoned through computational propaganda to induce harmful behaviors in large language models, even with data curation.
Researchers propose Decoupled Alignment, a training-free method using knowledge distillation to enhance LLM safety and prevent 'shadow alignment' during adaptation.
PERL is a new research model for Chinese Automatic Speech Recognition (ASR) error correction, addressing phonetic errors and length constraints.
Research explores using language models (LMs) as evaluators for other LMs, enhancing evaluation quality by increasing compute during test-time reasoning.
Researchers propose "Verbalized Sampling" to mitigate mode collapse in LLMs, attributing it to annotator bias favoring familiar text.
Research investigates LLM capability to translate conceptual research ideas into structured plans, demonstrating potential for scientific discovery acceleration.
Step-Tagging is a lightweight sentence-classifier framework designed to control and optimize Language Reasoning Model generation by reducing inefficient verification steps.
Research questions the consistent gains of tool use in web agents, citing limited experimental scales and non-comparable settings in prior studies.
FlowBot proposes using bilevel optimization and textual gradients to automatically induce LLM workflows, moving beyond human-crafted pipelines.
Researchers propose new algorithms to segment human-LLM co-authored text, moving beyond binary classification to localize generated segments.
Researchers propose MemTrace, a novel framework for tracing and attributing errors in large language model memory systems to improve reliability.
A research paper introduces ArogyaSutra, a multi-agent framework for multimodal medical reasoning in Indic languages, targeting low-resource healthcare.
Research proposes LLM agents can self-manage context via 'state proprioception' to overcome long-horizon context window limitations.
Research introduces a benchmark for evaluating AI safety against instruction conflict, embedded commands, and policy ambiguity, moving beyond simple pass/fail metrics.
Research describes a system combining speaker diarization with a Qwen3-ASR-1.7B model for multilingual two-speaker conversational speech.
Research introduces persona vectors to systematically audit open-weight LLMs across 53 behavioral traits, revealing suppressed or resistant behaviors.
Research explores augmenting natural language interfaces for data exploration with semantic and context-aware query recommendations over multi-table relational databases.
Research from arXiv explores how large language models, as a new communication medium, may transmit cultural patterns into human language.
Research explores integrating machine unlearning with RAG to manage sensitive information and harmful content retention in LLMs.
L-MARS is an open multi-agent legal QA system that uses agentic search and judge-driven evidence checks to audit citation faithfulness, addressing a common LLM failure in legal applications.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion