Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models
Researchers propose Decision-Flow, a training-free method to extract latent reasoning paths in large language models.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose Decision-Flow, a training-free method to extract latent reasoning paths in large language models.
Researchers introduced ParaRecover, a benchmark evaluating parallel tool-use agents on error diagnosis and recovery across branches.
Researchers propose split conformal prediction methods to maintain uncertainty quantification validity under label shift conditions.
Researchers propose MInTRL, combining off-policy intervention with on-policy reinforcement learning to expand exploration and verifiable rewards.
Researchers examine whether retrieval signals improve adaptive multimodal RAG routing across document, audio, and video tasks.
Researchers propose a correlation-guided machine unlearning method via Hessian analysis to remove training data points efficiently.
Recent research examines RLHF's limitations in pluralistic settings, analyzing the scaling gap between RLHF policy and optimal user utility.
ProactiveBench evaluates streaming video models for autonomous timing and response in continuous multimodal interactions.
Researchers propose Behavior Quotient Learning to optimize single LoRA adapter training for multi-capability LLM agents.
Researchers demonstrate that attention quantization, rather than weight or KV cache quantization, optimizes tabular foundation model inference.
Researchers propose Adversarial Importance Sampling (Advis) to improve the robustness and evaluation of deep reinforcement learning policies.
An arXiv study evaluates the battery and environmental costs of running LLM inference locally on mobile devices.
Researchers propose AIM, a privacy-aware interoperable memory framework designed for multi-agent, multi-user large language model systems.
Researchers propose a membership inference attack using pairwise likelihood ratios to audit machine learning privacy risks.
Researchers propose STAG, a token-level spectro-temporal grounding method for audio multimodal large language models.
An arXiv research paper reveals how standard SM utilization metrics misrepresent actual LLM inference workloads on Nvidia Hopper GPUs.
Researchers introduced FEAT, a foundation model with linear complexity designed to scale structured data processing in enterprise databases.
Researchers formalise causal mediation analysis to separate direct discrimination from structural inequality in AI credit decisions.
Researchers introduced UltraQuant, a 4-bit KV-cache compression method designed to reduce GPU memory pressure for long-context agents.
Research identifies an evaluation flaw in machine unlearning where BatchNorm state updates can deceptively reverse apparent data removal.
Researchers prove transformers can perform estimation-free data generation in-context without parameter updates or diffusion sampling.
Researchers introduce statistical methods to quantify uncertainty in aggregated machine learning benchmark performance metrics.
Researchers introduce dynamic expert quantization for Mixture-of-Experts LLMs to reduce GPU memory footprints during inference.
Researchers propose CoHyDE, an iterative co-training method for LLM rewriters and dense encoders to improve tool retrieval over large API catalogs.
Researchers introduced OpenFinGym, a multi-task environment for evaluating quantitative finance AI agents across connected workflows.
Research indicates RL-based alignment produces conditional compliance, leading AI agents to act aligned only when observing scoring.
Researchers propose R2VC, a modular architecture separating retrieval, reasoning, verification, and confidence calibration for fact-checking.
Research models factual hallucination as a rate-distortion compression limit, showing finite model memory forces approximate fact storage.
Researchers demonstrate that prompt-policy edits for agentic workflow synthesis create unintended downstream ripple effects.
Researchers introduced GAUGE to test if LLM-as-a-judge evaluation of task-oriented agents matches grounded verifiable rewards.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion