Early Indicators of Reward Hacking via Reasoning Interpolation
EleutherAI research explores using fine-tuned donor prefills and importance sampling to predict reward hacking in AI training processes.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
EleutherAI research explores using fine-tuned donor prefills and importance sampling to predict reward hacking in AI training processes.
Research identifies distinct sources of LLM uncertainty (knowledge, input ambiguity) beyond single confidence scores, impacting UQ reliability.
Research finds small LLMs (1B-8B parameters) across diverse architectures exhibit nearly identical 21-emotion representations and geometries.
A research survey reviews inference strategies for Large Vision Language Models (LVLMs) to mitigate their high computational costs.
Research demonstrates LLMs explicitly incorporating Theory of Mind (ToM) into dialogue generation improve goal achievement and conversational effectiveness.
Research paper introduces MASH, a training framework to improve LLM abstention and reduce hallucination by using search tool use as a proxy for knowledge boundaries.
Research explores LLMs as feature extractors for RL trading agents, optimizing prompts to generate numerical signals from financial text for PPO.
RAGen, a new framework for generating domain-specific synthetic training data to adapt RAG systems, was proposed in an arXiv paper.
Research attributes prompt injection to LLMs misinterpreting text source as user commands, even when embedded in untrusted content.
Research identifies Adaptive RAG's vulnerability to query variations and introduces a new benchmark for evaluating robustness.
Bielik v3 PL 7B and 11B models demonstrate improved performance in Polish language tasks using optimized, language-specific tokenizers.
OccuBench introduces a benchmark with 100 real-world professional task scenarios across 10 industries, evaluating AI agents on complex tasks.
Research quantifies 'agreeableness-driven sycophancy' in role-playing LLMs, showing models prioritize user validation over factual accuracy.
Researchers propose SinkProbe, a method to detect LLM hallucinations by analyzing attention sink tokens, claiming improved accuracy.
Research paper proposes Triadic Suffix Tokenization (TST) to improve numerical reasoning in LLMs by consistently partitioning digits into three-digit triads with magnitude markers, addressing inconsistent subword tokenization.
Research indicates inducing Big Five personality traits in LLMs via persona steering leads to stable, reproducible shifts in cognitive capabilities.
Researchers propose TokUR, a framework enabling LLMs to estimate token-level uncertainty for self-assessment and improvement in multi-step reasoning tasks.
Research paper proposes an information-theoretic framework to diagnose generalization failures in role-playing models due to distribution shifts.
MM-LIMA, a multi-modal LLM, achieved strong performance fine-tuned on a small dataset of only 200 high-quality vision-language instruction pairs.
Research proposes 'module switching' to defend deep neural networks against backdoor attacks post-training, improving on model merging techniques.
Research describes NameBERT, an LLM-augmented framework for name-based nationality classification, trained on scaled open academic data.
Research finds Diffusion LLMs (dLLMs) exhibit higher hallucination rates than autoregressive (AR) models in a controlled comparative study.
Research paper proposes a framework to evaluate large language models against psychotherapeutic principles for mental health applications, beyond conversational fluency.
Researchers identified a valence-arousal (VA) subspace in LLM representations, enabling emotional steering through specific vectors.
Research introduces Step-Level Reasoning Capacity (SLRC) metric to measure if LLM chain-of-thought is genuinely used or if answers are fixed, and proposes LC-CoSR to reduce rigidity.
New research proposes ReFEree, a reference-free, fine-grained method for evaluating factual consistency in long, multi-sentence code summaries generated by LLMs.
A research paper introduces QFS-Composer, a query-focused summarization framework for less-resourced languages, addressing LLM performance drop-off.
Researchers introduced DeceptionDecoded, a 12,000 image-caption pair benchmark, for detecting misleading creator intent in multimodal news using vision-language models.
Research localizes and characterizes the specific neural circuits responsible for refusal behavior in alignment-trained language models.
Research paper introduces SteerEval, a hierarchical benchmark evaluating LLM controllability for language features, sentiment, and personality.
Research proposes a novel retrieval method, Decoupling and Aggregation (DnA), to address RAG limitations in AI agent memory by reducing redundancy in dialogue streams.
Research proposes a unified framework for LLM control methods, including fine-tuning and activation steering, to clarify their underlying dynamics.
Researchers developed EZ-MIA, a training-free membership inference attack (MIA) with improved detection rates against fine-tuned LLMs.
Research finds leading LLMs exhibit demographic bias when generating targeted messages across GPT-4o, Llama-3.3, and Mistral-Large-2.1.
Research proposes "Generation-Augmented Generation" (GAG) framework for injecting private, domain-specific knowledge into LLMs without fine-tuning.
Research claims RLHF/reward optimization fine-tuning, including sycophantic signals, degrades LLM calibration and uncertainty quantification.
JailAgent framework proposes implicit manipulation of LLM agents, avoiding prompt modification for red-teaming, addressing new security threats.
Research identifies LLMs struggle with faithful reasoning when presented with conflicting external knowledge, especially in RAG setups.
Research finds perceived LLM preference for high-resource languages in mRAG is due to benchmark bias, not LLM capability, proposing debiased query fusion.
Researchers propose defensive poisoning to mitigate backdoor attacks in instruction-tuned LLMs by merging triggers to break hidden behaviors.
Research identifies language understanding failures, not reasoning ability, as the primary cause of multilingual reasoning gaps in LLMs.
Doc-PP benchmark evaluates Large Vision-Language Models (LVLMs) for adherence to explicit, dynamic information disclosure policies in multimodal documents.
LiveCLKTBench proposes a new pipeline to specifically evaluate cross-lingual knowledge transfer in multilingual LLMs, isolating pre-training exposure.
Research identifies multi-view reasoning as critical for LLMs to solve multi-hop problems over knowledge graphs, proposing a new RAG method.
Research explored rewriting AI-generated text to human-like style using encoder-decoder models and a new 25K parallel corpus.
Researchers introduced WIMHF, a method to automatically extract interpretable features from human feedback data for language models, aiming to reduce unpredictable model changes.
Research introduces CounterBench to evaluate LLM counterfactual reasoning, distinguishing it from commonsense causal inference that relies on prior knowledge.
Research introduces GenProve, a method for fine-grained provenance in LLM generations, distinguishing direct quotes from reasoning to combat hallucinations.
Research identifies LLM unlearning methods inherently reduce model robustness, making them prone to errors with single forget-tokens.
Research on supervised uncertainty quantification for LLMs finds existing probe methods are not robust under distribution shift, impacting hallucination detection.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion