The Randomness Floor: Measuring Intrinsic Non-Randomness in Language Model Token Distributions
Research introduces Entropic Deviation (ED) to measure intrinsic non-randomness in language model token distributions across various models and prompts.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces Entropic Deviation (ED) to measure intrinsic non-randomness in language model token distributions across various models and prompts.
Research finds hypernetwork-based LLM adaptation methods (e.g., Doc-to-LoRA) fail significantly (46.4% accuracy) when new facts contradict pretraining knowledge.
Research models how AI trading agents with similar market state representations can cause systemic instability in financial markets.
Research claims SFT-then-RL pipeline for LLM reasoning outperforms mixed-policy methods, attributing prior mixed-policy gains to a DeepSpeed optimizer bug.
Research identifies a 'backdoor mechanism' causing catastrophic overfitting in Fast Adversarial Training (FAT), leading to poor generalization in neural networks.
Research challenges the assumption that parameter-efficient fine-tuning (PEFT) methods like LoRA ensure memory efficiency in LLMs due to intermediate tensor scaling.
Research explores using machine learning to guide primal heuristics for Mixed Binary Quadratic Programs, aiming for faster, high-quality solutions.
Research identifies standard LLM evaluation metrics (confusion matrix) are misleading for imbalanced, cost-asymmetric tasks like literature screening.
Research identifies 'Self-Preference Bias' in LLM judges, where models favor their own outputs, impacting automated evaluation systems.
Research highlights that single-seed benchmarks for Bayesian deep learning in limited-data settings can misrepresent model stability due to high variance.
Research finds frontier LLMs exhibit 'heuristic collapse' when giving investment advice, failing to integrate full user context.
Research shows Bayesian deep learning model rankings are unstable and dataset-dependent, particularly with scarce data, challenging standard evaluation assumptions.
Research paper identifies failure modes in standard on-policy distillation (OPD) for LLMs and proposes fixes to improve learning signal stability.
Research questions the effectiveness and nature of Chain-of-Thought (CoT) reasoning in LLMs, attributing successes and failures to data distribution.
Research explores using dataset statistical effect size to predict model performance and determine data sample size sufficiency prior to training.
Research identifies 'secondary risks' in LLMs: non-adversarial, subtle failure modes during benign interactions, distinct from jailbreak attacks.
Research identifies 'supernodes' in LLM feed-forward networks, where 1% of channels account for nearly 60% of loss sensitivity in Llama-3.1-8B.
New diffusion model sampling algorithms achieve exponential speedup (polylogarithmic steps) for high accuracy, improving prior methods.
Researchers propose probe-based data attribution to identify training datapoints responsible for specific LLM behaviors by analyzing activation differences.
Research based on 33 interviews with smart-home AI designers details current approaches to ethics and expectations management at Amazon, Microsoft, and Google.
Research indicates general Process Reward Models (PRMs) fail to detect silent errors and logical flaws in LLM-driven data analysis agents.
Clotho introduces a pre-generation test adequacy measure for LLM inputs, aiming to reduce human judgment reliance and post-inference testing.
Research identifies training-inference inconsistency in LLM-based recommender systems using supervised fine-tuning and beam search.
Researchers propose a method to improve machine learning model robustness by identifying and mitigating spurious correlations without group annotations.
Research introduces CoRT, a black-box multi-turn red-teaming framework to find concealed regulatory-violating risks in financial LLMs.
FedSLoP, a new federated optimization algorithm, uses low-rank gradient projections to improve convergence and reduce communication/memory costs in federated learning.
AgenticCache, a new planning framework for embodied AI agents, reuses cached plans to significantly reduce LLM calls, improving latency and cost.
PyPOTS, an open-source Python ecosystem, introduces end-to-end data mining for partially-observed time series (POTS) with integrated missing-value handling.
RouteNLP is a research framework proposing closed-loop LLM routing to optimize cost by directing queries to different model sizes based on difficulty.
Research introduces model-agnostic explainers based on Shapley and Owen values for Temporal Graph Neural Networks (TGNNs) to improve transparency.
Research characterizes the impact of prompt and response characteristics on LLM inference energy costs, highlighting sustainability and financial feasibility.
MermaidSeqBench, a human-verified benchmark, has been introduced to evaluate LLM correctness for natural language to Mermaid sequence diagram generation.
LongFlow is a research technique to compress KV caches, reducing memory consumption and bandwidth pressure for LLMs generating long output sequences.
Research identifies and evaluates 'sycophancy' in LLMs within agentic financial tasks, where models prioritize agreement over correctness.
Research paper proposes GWT, a scalable optimizer state compression method for large language model training, reducing memory overheads.
Multi-agent LLM tutoring systems incur higher latency and cost due to compounded API calls compared to single-agent systems, per arXiv research.
Research identifies and quantifies the impact of 'spurious features' (implicit noise) in grounding data on RAG system robustness, proposing improvement methods.
A research survey explores split learning as a method for fine-tuning LLMs, addressing data privacy concerns and computational costs.
Research finds classical CPU-based algorithms consistently outperform GPU-based AI methods, including generative models and reinforcement learning, on the Maximum Independent Set problem.
Research proves that verifying quantized Graph Neural Networks (GNNs) with global readout is computationally intractable (coNEXPTIME-complete).
Research proposes Selective Conformal Risk Control (SCRC), a framework combining conformal prediction with selective classification for reliable uncertainty quantification.
Research revisits parameter sharing in LoRA fine-tuning, finding inner A matrices are highly similar across multiple LoRAs, suggesting efficiency gains.
MERIT, a modular framework using GPT-4o-mini, achieved 81.65% F1 on MMFakeBench for multimodal misinformation detection, outperforming GPT-4V.
Super-DeepG, a new method for formally verifying neural networks against geometric perturbations in image data, improves linear relaxation techniques.
Research explores using instrumental regression and GMM to address moral hazard in data-driven policy-making, where individual actions are unobserved.
Microsoft's GitHub is moving to metered billing for AI features, indicating a shift from fixed-cost 'all-you-can-eat' models to usage-based pricing.
Adaptive ultrasound imaging leveraging physics-informed AI, demonstrated on Hugging Face.
OpenAI detailed its safety framework for ChatGPT, including model safeguards, misuse detection, policy enforcement, and expert collaboration.
OpenAI models (GPT, Codex) and Managed Agents are now available on AWS, enabling enterprises to build AI securely within their AWS environments.
Applied Intuition discusses deploying AI in highly adversarial physical environments across mining, drones, trucks, and warships.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion