TAPO: Transition-Aware Policy Optimization for LLM Agents
Researchers introduce TAPO, a post-training method optimizing LLM agents using dense environmental feedback rather than sparse task rewards.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduce TAPO, a post-training method optimizing LLM agents using dense environmental feedback rather than sparse task rewards.
Researchers introduce ClawTrack, a dual-assessment benchmark designed to evaluate both the task outcomes and trace-level reasoning of autonomous agents.
Research shows mathematically equivalent expert aggregation orders in sparse MoE models like DeepSeek-V4-Flash cause divergent outputs.
Researchers propose an Information Bottleneck approach to improve the mathematical faithfulness of explanations in time series forecasting.
Researchers propose LEDGERMIND, a framework for evaluating multimodal agents using a structured evidence ledger for step-by-step verification.
Researchers propose Adaptive Anticipatory Policy Trees to eliminate autoregressive decoding delays in multimodal computer-use agents.
Research trains a chain-of-thought (CoT) reasoning-enabled LLM to classify cybersecurity detections, aiming to reduce alert fatigue in SOCs.
Research explores integrating causality into algorithmic recourse to ensure recommended changes genuinely improve qualifications, not just game classifiers.
Research identifies a structural issue in on-policy self-distillation (OPSD) for LLMs, proposing $\beta$-OPSD as a more stable policy optimization approach.
KAISEN proposes a five-phase, reproducible audit pipeline for clinical risk models to identify and diagnose error rate disparities across patient subgroups.
Researchers demonstrate that Earth observation foundation models improve regional forest biomass monitoring, addressing data sparsity in carbon tracking.
Researchers introduce Divergence Decoding, a training-free inference framework that fuses generalist reasoning and specialist domain models.
Academic paper proves that retraining neural networks on pooled datasets can cause unexpected decision reversals, undermining model reliability.
Research finds Diffusion Language Models (DLMs) are less robust to input noise and adversarial attacks than autoregressive (AR) models like Llama 3.
OneShot introduces a neural scoring approach for large-scale retrieval systems, improving ranking accuracy while maintaining indexing efficiency.
Research explores how fine-tuning LLMs on one strategic game impacts their performance on other games, using normal-form games as a testbed.
HealthCAT, an interpretable encoder-only Transformer, predicts health indicators from wearable sensor data, prioritizing interpretability over accuracy.
GyRot proposes a quantization framework and hardware accelerator to improve low-bit LLM inference by integrating rotation and fine-grained group quantization.
LightRot introduces a lightweight rotation scheme and hardware architecture for energy-efficient, accurate low-bit large language model inference.
Research presents a theoretical error analysis for Engression, a neural-network-based conditional distribution learning method using the energy score.
New research addresses the challenge of robustly estimating sparse numerical vectors under Local Differential Privacy (LDP) against poisoning attacks for multi-item users.
STEREODISCO framework measures and discovers previously unexamined stereotypical associations and biases in LLMs using semantic differential method.
Research paper introduces a safety-gated agentic supervisory control system for industrial processes, using an open-weight LLM with auditable rule-based constraints.
Research introduces Dynamic Spectral Filtering (DSF), an operator-centric approach for temporal graph learning, allowing propagation mechanisms to evolve.
Research demonstrates "sponge attacks" on Spiking Neural Networks (SNNs) can inflate energy consumption and synaptic workload, impacting efficiency.
Researchers introduced Echoverse, a framework for generating deep, evolving synthetic environments to train computer-use agents at scale.
Research explores Group-Reflective Self-Distillation for agentic reinforcement learning, aiming to improve training by refining sparse terminal rewards.
Researchers evaluate six deep learning weather emulators, showing superior performance over traditional dynamical forecasting models.
A research survey reviews methods for Uncertainty Quantification (UQ) in deep learning, focusing on ensemble and approximate Bayesian approaches.
ShadowDancer is a research approach enabling frame-level control of video world models through unified dynamics representations from video and its 'shadow'.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion