Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling
Research shows multi-layer linear probes improve detection of 'wrong' or deceptive LLM outputs, increasing AUROC by +29% on specific tasks.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research shows multi-layer linear probes improve detection of 'wrong' or deceptive LLM outputs, increasing AUROC by +29% on specific tasks.
Researchers propose a latent-space language steering method using PCA to reduce unintended code-switching in multilingual LLMs during inference.
Research systematically compares prompt design, generator models, and source data for synthesizing high-quality LLM pretraining data.
Research paper proposes numerically stable and federated power transforms, addressing existing instabilities in data preprocessing methods.
Open-weight models achieved IOI gold medal performance by scaling test-time compute, demonstrating advanced reasoning capabilities in programming.
Research evaluates LLMs against the Chomsky Hierarchy to assess formal reasoning capabilities, finding current benchmarks inadequate.
LongCoT introduces a new benchmark for evaluating long-horizon chain-of-thought reasoning in LLMs across various domains.
Research details 'model reprogramming' to perform membership inference attacks without shadow models, reducing computational cost.
Research introduces new neural architectures outperforming existing sequence-to-sequence models on synthetic benchmarks for reference resolution in code.
Research explores KL divergence for mixed-precision quantization in hybrid SSM-Transformer LLMs, aiming for efficient edge device deployment.
OpenAI introduces GPT-Rosalind, a frontier reasoning model for drug discovery, genomics, and scientific research workflows.
OpenAI launched 'Trusted Access for Cyber' program, providing security firms access to GPT-5.4-Cyber and API grants for cyber defense.
Google DeepMind's Gemini 3.1 Flash TTS introduces granular audio tags for expressive AI speech generation, offering precise control.
OpenAI updated its Agents SDK, adding native sandbox execution and a model-native harness for building secure, long-running AI agents.
HCompany introduced HoloTab, an AI browser companion for enhanced web interaction. Details on specific capabilities are limited.
Research indicates LLMs, including GPT-4o, struggle with abstract meaning comprehension beyond current expectations on the SemEval-2021 ReCAM task.
Research evaluates various PDF parsing and chunking methods for financial Q&A in RAG systems, highlighting challenges with heterogeneous content.
Research indicates LLMs struggle with reliable instruction following across nuanced, analogous prompts despite high benchmark scores on IFEval, impacting real-world performance.
Research paper surveys actionable mechanistic interpretability methods for LLMs, categorizing techniques for locating, steering, and improving model behavior.
Research introduces Block Diffusion Draft Trees for speculative decoding, improving LLM inference speed by generating draft blocks in a single pass.
New research proposes Filtered Reasoning Score to evaluate LLM reasoning quality independently of output accuracy, addressing flawed reasoning for correct answers.
Universal NER project released v2, an expanded multilingual Named Entity Recognition (NER) benchmark for evaluating LLMs across more languages.
Research shows simple lexical constraints (banning a single character or word) cause instruction-tuned LLMs to lose 14-48% comprehensiveness.
Research finds LLMs are severely overconfident (ECE 0.35-0.64) on tabular question answering, significantly worse than textual QA (0.10-0.15).
Research trains LLMs to perform human-like, meaning-preserving edits of inappropriate argumentation using reinforcement learning.
Research proposes an Item Response Theory (IRT) framework for extensible LLM benchmarking, calibrating new benchmarks to existing suites using anchor items.
New benchmark, GlotOCR Bench, shows current OCR models struggle with generalization across 100+ Unicode scripts, performing poorly on low-resource languages.
Researchers propose "cooperative paging" to manage long LLM conversations: evicted content is replaced with keyword bookmarks, and the model can recall full text.
Research introduces MulTypo, a multilingual typo generation algorithm, to evaluate LLM robustness against human-like typographical errors in diverse languages.
Research paper introduces CodeSpecBench, a new benchmark for evaluating LLMs' ability to generate executable behavioral specifications (pre/postconditions) from natural language.
Research demonstrates a method to compile activation steering into LLM weights, creating stealthy backdoors that trigger jailbreaks under specific inputs.
Research introduces CompliBench, a benchmark for evaluating LLM judges' ability to detect compliance violations in dialogue systems.
Research identifies visual token dominance as the core bottleneck in large Vision-Language Model (LVLM) inference efficiency, proposing a taxonomy of techniques.
New arXiv paper proposes benchmarks for Large Vision-Language Models (LVLMs) to test deflection and hallucination with conflicting visual and textual evidence.
AlphaEval proposes a new framework for evaluating AI agents in production environments, accounting for heterogeneous, multi-modal inputs and implicit constraints.
Research proposes 'reasoning calibration' to improve LLM factuality in long-form generation by enabling models to estimate reliability of claims.
Research finds LLMs exhibit the 'Identifiable Victim Effect,' prioritizing narratively described individuals over statistically larger groups in resource allocation.
ReasonXL paper claims LLMs can be fine-tuned to reason in non-English languages without performance loss, addressing English-centric reasoning.
Researchers propose ToxiTrace, a BERT-style model method using LLM guidance for explainable Chinese toxic content detection with fine-grained toxic span identification.
Research proposes a novel method, GRADE, using gradient subspace dynamics to probe LLM internal knowledge gaps, aiming for better confidence detection.
Research introduces Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework to improve LLM precision and reproducibility in text categorization.
Research proposes multi-agent, multi-format approach for LLMs to understand complex spreadsheets, addressing layout cues and scale limits.
Research explores using LLMs to evaluate data privacy and AI safety in contexts with imperfect information, moving beyond complete context assumptions.
Research explores if LLMs possess 'privileged knowledge' about their own answer correctness from internal states, beyond external observation.
Research paper proposes "SeedPrints" method to identify the random seed used to train a Large Language Model for provenance and attribution.
New arXiv research questions if VLMs genuinely understand candlestick charts for stock forecasting, citing inadequate benchmarks.
Researchers demonstrated a clean-label backdoor attack on Graph Neural Networks (GNNs), manipulating predictions without altering training node labels.
Research paper proposes Parcae, a new training recipe for stable, looped language models that scales quality via recurrent computation within fixed parameters.
Research explores Monte Carlo Stochastic Depth (MCSD) to enhance uncertainty quantification (UQ) in deep learning, building on MC Dropout methods.
Research proposes provably replicable reinforcement learning algorithms with linear function approximation to address experimental variability.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion