Learning When to Reason for Text-to-SQL via SFT and DPO
New research proposes AutoThinkSQL, a framework that dynamically decides when to apply reasoning chains for Text-to-SQL queries, reducing inference overhead.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
New research proposes AutoThinkSQL, a framework that dynamically decides when to apply reasoning chains for Text-to-SQL queries, reducing inference overhead.
New research introduces LENS, a protocol for evaluating how well machine unlearning algorithms suppress disinformation-aligned narratives in large language models.
Research details a simple language normalization method for cross-lingual speaker verification, improving performance on the TidyVoice 2026 Challenge.
Research applies a two-stage LLM pipeline (Gemini 2.5 Pro, Gemini 2.5 Flash) to detect documentation inconsistencies in 3,000 health records.
Researchers developed ADAGE, a language-agnostic pipeline combining native-speaker curation with LLM-assisted generation for culturally-grounded analogical reasoning benchmarks.
New research proposes attention-guided strategies for contrastive decoding (e.g., DoLa) in LLMs to improve factuality beyond vocabulary divergence.
Researchers used Low-Rank Adaptation (LoRA) fine-tuning to achieve 80% accuracy in gender-inclusive text rewriting for the LT-EDI 2026 Shared Task.
Research explores training-free methods to personalize language model toxicity sensitivity during inference, moving beyond global alignment for subjective contexts.
Researchers introduced IndicTalk, a large persona-based multilingual conversational corpus for Indic languages, addressing scarcity of high-quality code-mixed dialogue data.
Research introduces a method for multi-hop question answering by jointly evolving graph and text memories, enabling coordination of relational and textual evidence.
Research compares BERT-based models and LLMs for Named Entity Recognition in Marathi, a low-resource language, finding uncertain effectiveness for LLMs.
Researchers introduced JOLT, a method to jointly optimize subword vocabularies for greedy longest-match tokenization, improving compression for LLMs.
Research finds Activation Oracles (AOs), LMs explaining other models' internals, can develop concept-specific blind spots due to training data.
LA-RL introduces a label-aware self-reflection method for reinforcement learning in information extraction, improving structured output correction.
Research evaluates LLM understanding and generation of novel Chinese xiehouyu riddles using MCQs and free-form explanation generation to test reasoning.
Research indicates LLM debate patterns differ across languages, with Chinese models less prone to repeating arguments compared to other languages tested.
Research finds fine-tuning small models with legal context improves their accuracy on legal Q&A, even with retrieved law.
Research identifies that offensive language detection models degrade across datasets and languages, proposing a framework to diagnose and optimize cross-domain generalization.
Research indicates that diagrams, particularly Euler diagrams, can improve LLM syllogistic reasoning performance compared to natural language or logical notation.
Research paper proposes a novel method for dynamic evaluation of multimodal automated fact-checking systems to avoid contamination from outdated claims.
Research explores why Joint-Embedding Predictive Architectures (JEPAs) are not standard for text encoders, citing a mismatch with language's conditional structure.
Researchers developed formally verified IEEE-754 FP32 and BF16 arithmetic for ARCH HDL, a language intended for LLM generation, ensuring mathematical correctness.
Research proposes a 'frozen model' architecture with a persistent memory of verified solutions for 100% accuracy and zero inference tokens.
Earnings25, a new 500-hour finance-domain speech benchmark for ASR evaluation on English earnings calls, is introduced by arXiv research.
Research finds LLM prompt tone significantly impacts inference cost via output token length, with less effect on accuracy across seven tones on MMLU.
Research finds appending a two-word confirmation tag to a question changes LLM responses, measuring this 'tag effect' across 45 models.
Researchers propose Mixture of Language Group Experts (MoLGE) to improve performance and efficiency in massively multilingual automatic speech recognition models, addressing 'curse of multilinguality'.
Research explores pointer-augmented autoregressive generation for patent claims, addressing hierarchical constraints in structured text with LLMs.
Research compares LLM-based vs. lexicon-based sentiment analysis for detecting tail-risk signals from Reddit data on meme stocks like GME and AMC.
Research identifies two distinct failure points in compressed short-text generation: information loss in the codec or weak codes from the latent generator.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion