Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Re-FORC proposes an adaptive reward prediction method for Chain-of-Thought reasoning to enable early stopping and reduce compute costs by up to 26%.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Re-FORC proposes an adaptive reward prediction method for Chain-of-Thought reasoning to enable early stopping and reduce compute costs by up to 26%.
Researchers introduced Atlas 2, Atlas 2-B, and Atlas 2-S, three pathology vision foundation models addressing performance, robustness, and computational requirements for clinical deployment.
Research explored multi-agent Q-learning coordination in a tabular predator-prey gridworld, isolating coordination structure from approximation.
Research finds the pretraining domain, not the training objective, is the primary factor in differentially private medical imaging models' utility-privacy trade-off.
A research paper introduces Predictive Query Language (PQL), a domain-specific language for predictive modeling directly on relational databases.
Research introduces 'alternation metrics' for multi-agent systems to evaluate temporal fairness in resource access beyond aggregate payoffs, using the Honey-Jar Game.
Research explores Minimum Norm Interpolation (MNI) framework in overparameterized models, focusing on generalization under 2-uniform convexity.
Research introduces Math Education Digital Shadows (MEDS), a dataset to evaluate 14 LLMs' mathematical performance and biases across personifications.
Research finds Transformer models' high performance in intrusion detection is often due to flawed evaluation, not true temporal gains.
Research paper introduces a framework to analyze representation costs from parameter-space regularizers in data-fitting methods, including DNNs.
Research finds multimodal LLMs are vulnerable to stylistic jailbreak attacks beyond content-based methods, exploiting comprehension inconsistencies.
Researchers propose a consensus-based framework for evaluating large language models by measuring relative preference in complex, subjective tasks.
Researchers developed Humanly, a configurable environment to track and trace human-AI collaborative writing processes for audit and evaluation.
Research probes Qwen2.5-7B's internal representation of Colombian identity and socioeconomic status from linguistic cues using Natural Language Autoencoders.
Researchers introduced Copyright-Bench, a new benchmark to evaluate LLM agents' compliance with copyright law when reproducing external content.
Research explores baking documents into Gemma-4-e4b model weights via LoRA for closed-book QA, finding data quality critical over capacity.
Researchers conducted the first systematic study of faithfulness in document-grounded, multi-speaker podcast generation using LLMs, addressing ungrounded information.
New research proposes 'Ground Truth First,' a novel methodology for evaluating LLM agent memory by simulating life-scripts with ground truth facts before text generation.
Researchers propose MoE^2-LoRA, a new parameter-efficient fine-tuning (PEFT) method designed specifically for Mixture-of-Experts (MoE) large language models.
J-CoT proposes 'J-Space' as a method to improve chain-of-thought prompting in LLMs by recurrently propagating dense hidden vectors.
Research analyzes how different LLM architectures represent self-harm content, aiming to improve detection and governance of high-stakes safety issues.
A new research paper, DWT-Fusion, proposes a training-free, signal-based framework for detecting LLM-generated text, focusing on local and multiscale variations.
Research explores using reinforcement learning (RL) to mitigate task conflicts when merging multiple fine-tuned LLMs into a single model.
Researchers developed and validated the Spanish version of the Large Language Models Dependency Scale (LLM-D12-SP) to assess psychological dependency.
New research explores native multimodal pre-training from scratch to overcome limitations of text-only models and enhance real-world perception.
Researchers introduced MEUSLI, an open-science multilingual projector for connecting speech encoders (like Whisper) with large language models.
A new research paper proposes a multi-layer taxonomy of 14 capability domains and 91 subskills for large language model evaluation.
Research evaluates large language models (LLMs) as creativity evaluators, examining convergence and divergence with human judgments across six LLMs.
New research explores 'Skill Self-Play' for LLMs, enhancing capabilities through co-evolving skills and addressing the task diversity versus verification reliability dilemma.
Research surveys methods to translate EU AI Act obligations into testable, auditable requirements and verifiable evidence, exploring LLM-based agentic tools.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion