Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Re-FORC proposes an adaptive reward prediction method for Chain-of-Thought reasoning to enable early stopping and reduce compute costs by up to 26%.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Re-FORC proposes an adaptive reward prediction method for Chain-of-Thought reasoning to enable early stopping and reduce compute costs by up to 26%.
Research identifies explicit iteration complexity for exact data-driven inverse optimization of Integer Linear Programs using gradient-based methods.
Research proves ReLU networks with two hidden layers can exactly represent the maximum of up to 10 real numbers using rational linear algebra.
Research proposes General Value Functions (GVF) for remaining useful life (RUL) and failure-mode prediction in predictive maintenance.
Hopformer, a two-stage Transformer framework, is introduced for multi-variate time series forecasting by separating common trends from specific information.
Researchers propose SEM-DNN, a heteroscedastic neural simultaneous-equation estimator, to learn bidirectional causal interactions from observational data.
Research improves differentially private stochastic gradient descent (DP-SGD) accuracy by correlating privacy noise across iterations using model curvature.
A research paper proposes Deep Convolutional Large-Margin $\ell_p$-SVDD for visual anomaly detection, combining deep features with explicit margin-aware boundaries.
Research proposes a theory of indecisions for selective hypothesis testing to minimize abstention rates while maintaining target accuracy in high-risk scenarios.
Researchers propose Evaluation-as-a-Service (EaaS), a cloud-native microservices architecture for scalable AI monitoring with conformal guarantees.
Research explores making large text-to-image diffusion models more interpretable and manipulable for creative uses, focusing on interactive explainability.
Research explores agentic AI for automated, evidence-grounded root cause analysis of industrial anomalies, addressing explainability and data scarcity.
HiKV proposes a novel algorithm-hardware co-design to compress the KV cache in LLM decoding, tackling memory bottlenecks for long-context models.
Research explores Minimum Norm Interpolation (MNI) framework in overparameterized models, focusing on generalization under 2-uniform convexity.
Research proposes a convex optimization framework to generate theoretical correlation matrices with graph-based sparsity patterns, improving matrix completion.
Research proposes TRACE-ROUTER, a new routing mechanism for agentic AI applications that optimizes LLM selection based on long-horizon, task-level outcomes.
Research proposes treating LLM prompts as a first-class data type within database management systems for better optimization and governance.
Research introduces Simulation-Based Empirical Bayes (SBEB) for simultaneous inference when likelihoods are available only through simulators.
Research introduces SurvDiff, a diffusion model designed for generating synthetic survival data, addressing challenges of incomplete event information.
CausalForge is a framework for automated theoretical research in causal inference, designed to address the unreliability of LLM-based reviewers.
Research explores meta-learning for speaker-dependent voice fatigue models to improve performance and efficiency over traditional mixed-effect models.
Research proposes a Decentralized Multi-Agent Swarm (DMAS) architecture using autonomous agents for security in Industrial IoT (IIoT) environments.
Research proposes a layer-wise LoRA fine-tuning method using a similarity metric to improve LLM predictive performance efficiently.
Researchers introduced Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models addressing performance, robustness, and computational requirements for clinical deployment.
Research explores differentially private federated learning for imbalanced clinical data using SMOTETomek and FedProx to balance privacy and utility.
Research tests sparse attention mechanisms, finding attention patterns do not reliably indicate which parts of context are used by large language models for answers.
Research finds multimodal LLMs are vulnerable to stylistic jailbreak attacks beyond content-based methods, exploiting comprehension inconsistencies.
Researchers developed Humanly, a configurable environment to track and trace human-AI collaborative writing processes for audit and evaluation.
Research probes Qwen2.5-7B's internal representation of Colombian identity and socioeconomic status from linguistic cues using Natural Language Autoencoders.
MetaEvolve is a research framework enabling LLMs to develop meta-skills like self-reflection through reinforcement learning, improving test-time performance.
VLMs, when used for document understanding, may rewrite imperfect text into a more plausible form rather than faithfully transcribing it, as revealed by a new multilingual perturbation benchmark.
Researchers developed and validated the Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP) to assess psychological dependency.
Researchers introduced MEUSLI, an open-science multilingual projector for connecting speech encoders (like Whisper) with large language models.
Research surveys methods to translate EU AI Act obligations into testable, auditable requirements and verifiable evidence, exploring LLM-based agentic tools.
Researchers propose a consensus-based framework for evaluating large language models by measuring relative preference in complex, subjective tasks.
Research evaluates large language models (LLMs) as creativity evaluators, examining convergence and divergence with human judgments across six LLMs.
A new research paper proposes a multi-layer taxonomy of 14 capability domains and 91 subskills for large language model evaluation.
Researchers introduced Copyright-Bench, a new benchmark to evaluate LLM agents' compliance with copyright law when reproducing external content.
Research explores using reinforcement learning (RL) to mitigate task conflicts when merging multiple fine-tuned LLMs into a single model.
Research explores baking documents into Gemma-4-e4b model weights via LoRA for closed-book QA, finding data quality critical over capacity.
New research explores 'Skill Self-Play' for LLMs, enhancing capabilities through co-evolving skills and addressing the task diversity versus verification reliability dilemma.
A new 3B parameter model, Nanbeige4.2-3B, claims strong agentic capabilities and competitive reasoning across multiple domains, trained on 28T tokens.
A research paper benchmarks an open-weight 31B multimodal model (Gemma 4) using fine-tuning and RAG on the U.S. NRC Reactor Operator licensing exam.
Research finds small open-weight VLMs, Qwen2-VL-2B-Instruct and SmolVLM-Instruct, have internal uncertainty but struggle to express it under image degradation.
Researchers propose MoE^2-LoRA, a new parameter-efficient fine-tuning (PEFT) method designed specifically for Mixture-of-Experts (MoE) large language models.
New research introduces InteractComp, a benchmark for evaluating search agents on ambiguous queries requiring interactive clarification, addressing a common failure mode.
Research finds commercial LLMs, particularly Grok, vary in stability and transparency when evaluating ethnonationalist pseudo-science via API vs. web interfaces.
Research paper proposes a four-layer technical architecture for large model inference optimization, focusing on token-oriented techniques.
Researchers propose SURE-RAG, a method for Retrieval-Augmented Generation that verifies evidence sufficiency and uncertainty, addressing cases where retrieved passages are relevant but insufficient for an answer.
Research finds LLM interventions can improve cross-partisan receptivity to news but LLMs overestimate their own debiasing effectiveness in trials.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion