Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
Loong introduces a method for synthesizing long chain-of-thoughts using verifiers to improve LLM reasoning, particularly in domains like mathematics and programming, extending RLVR.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Loong introduces a method for synthesizing long chain-of-thoughts using verifiers to improve LLM reasoning, particularly in domains like mathematics and programming, extending RLVR.
Research introduces a physics-encoded inverse modeling approach for estimating Arctic snow depth from sparse, indirect observations.
Research proposes Shift-Aware Calibration (SAC) for fine-tuned Vision-Language Models like CLIP, improving confidence-accuracy alignment on unseen data.
Research finds LLM-generated authentication code has security vulnerabilities, even with iterative reprompting, when assessed against NIST SP 800-63B.
Research paper proposes an offline-to-online workflow using generative models and adaptive testing for ad creative optimization, leveraging historical A/B test data.
ANSR-DT is a neuro-symbolic framework for digital twins, integrating temporal anomaly detection, symbolic reasoning, and RL-based decision support.
Researchers propose a two-level Hierarchical Reinforcement Learning (HRL) framework using Soft Actor-Critic for sparse-reward, long-horizon tasks.
Researchers introduce CORVUS, a framework that optimizes LLM coding agent context by synchronizing changing files instead of appending static snapshots.
Research proposes a LoRA-based approach for Domain Incremental Learning to prevent catastrophic forgetting by consolidating shared knowledge across tasks.
Researchers propose ESRVS, a semi-supervised method for retinal vessel segmentation requiring only one annotated image and unlabeled data.
Research demonstrates integrating GNSS Zenith Wet Delay into AI weather models improves precipitation forecasts, addressing known underestimation.
Research explores a transformer-based model for automatically identifying conflicting and duplicate software requirements to enhance development efficiency.
Research finds that restricting the feasible set in constrained stochastic optimization can paradoxically increase statistical risk for projection estimators.
Research introduces an asynchronous, event-driven clustering algorithm for real-time detection of small event clusters in event camera data.
Research paper settles the problem of learning optimal linear contracts from data in an offline setting, showing Empirical Utility Maximization provides an ε-approximation.
Research proposes a physiology-guided self-supervised learning method for screening Aortic Valve Disease using photoplethysmography (PPG) signals.
Research identifies legal challenges and shortcomings in EU AI Act provisions for robustness and cybersecurity in high-risk AI systems.
Research introduces 'self-distillation of hidden layers' for self-supervised learning, aiming for efficient, stable high-level embedding generation.
Research proposes physics-constrained neural networks with embedded gradient networks for dynamic modeling of synchronous machines, ensuring energy-balance.
LanteRn is a research model improving visual reasoning for LMMs by using latent visual representations instead of verbalizing perceptual content.
Researchers propose optimizing Transformer inference on FPGAs for real-time anomaly detection in high-frequency financial time series.
Research proposes MTSF-ANO, a hybrid quantum-classical model using adaptive non-local observables for multivariate time series forecasting.
Researchers developed formally verified IEEE-754 FP32 and BF16 arithmetic for ARCH HDL, a language intended for LLM generation, ensuring mathematical correctness.
Research identifies two distinct failure points in compressed short-text generation: information loss in the codec or weak codes from the latent generator.
New research proposes Direct Corpus Interaction (DCI) for agentic search, guiding fine-grained corpus exploration with relevance estimates.
Researchers propose SINT-Flow, an LLM-based framework with five operators for fully automated, end-to-end schema integration across diverse input tables.
New research introduces LENS, a protocol for evaluating how well machine unlearning algorithms suppress disinformation-aligned narratives in large language models.
Researchers propose Zhijing, a framework for measuring and integrating social intelligence in LLMs, including the SoMBench benchmark.
Researchers introduced LEDOM, a purely right-to-left autoregressive language model (2B/7B parameters), finding distinct abductive reasoning capabilities.
New research explores if transformers can dynamically adapt their reasoning strategies (latent algorithm routing) based on input data characteristics.
Research compares LLM-based vs. lexicon-based sentiment analysis for detecting tail-risk signals from Reddit data on meme stocks like GME and AMC.
Omni-Prune proposes query-aware token pruning for omnimodal LLMs to reduce inference latency and GPU memory for audio-video inputs.
Research finds LLM prompt tone significantly impacts inference cost via output token length, with less effect on accuracy across seven tones on MMLU.
Earnings25, a new 500-hour finance-domain speech benchmark for ASR evaluation on English earnings calls, is introduced by arXiv research.
Researchers propose Mixture of Language Group Experts (MoLGE) to improve performance and efficiency in massively multilingual automatic speech recognition models, addressing 'curse of multilinguality'.
Research introduces Cognitive Attribution Graphs (CAGE) to improve inline citation generation in long-form LLM outputs, addressing 'attribution ambiguity'.
Researchers propose using top-k log probabilities, a readily available inference signal, to monitor LLM performance and prioritize interventions.
Research quantifies the 'tokenizer tax' for Indian languages, showing LLMs incur higher processing costs due to English-centric subword tokenizers.
Research proposes a 'frozen model' architecture with a persistent memory of verified solutions for 100% accuracy and zero inference tokens.
SyRuP, a new method, aims to improve LLM adherence to complex system prompts during decoding without requiring model tuning or reranking.
Research identifies a 'blind spot' in AI agent long-term memory systems where retrievers fail to link implicit knowledge to queries.
Research evaluates closed-loop validation-repair for clinical LLMs in healthcare to achieve structured output schema compliance (ICD-10, CPT, HL7 FHIR).
INS-ActBench, a new benchmark, evaluates LLMs on professional actuarial tasks requiring auditable, context-grounded, and tool-executable decisions.
Research indicates self-improving agents using self-authored tests for verification can show high internal scores while real performance degrades.
New research from arXiv proposes Cross-Attention Calibrated Deduplication to improve RAG system efficiency by identifying and removing redundant data chunks.
Researchers introduced LEX-EC, a black-box audit framework for zero-shot LLM personality classification, using lexical ablation to distinguish signal from marginal-distribution effects.
CONSISTRE is a new framework designed to improve consistency in document-level relation extraction using large language models by enforcing relational constraints.
Researchers propose "masked distillation" to internalize Chain-of-Thought reasoning in language models, reducing inference latency and cost.
Research explores Parallel Autoregressive Decoding (PARD) for block diffusion language models, showing better alignment with left-to-right generation.
New research introduces a two-layer evaluation framework to separate language model execution evidence from correctness, exposing how models fail beyond accuracy scores.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion