Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Researchers propose Entropy-Guided Supertokens to compress LLM reasoning traces, splitting tokens into structural and organic types.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose Entropy-Guided Supertokens to compress LLM reasoning traces, splitting tokens into structural and organic types.
Researchers propose Policy-Masked Private Experts to restrict forward-pass routing to authorized MoE parameters based on policy.
Researchers demonstrate that Rectified Flow generative models leak membership signals of training data along their interpolation paths.
Research shows LLM-as-a-judge evaluation accuracy degrades when evidence extraction is separated from final verdict generation.
Researchers introduce Leak It, a probabilistic method to extract training data from black-box LLMs by analyzing output sample distributions.
Researchers propose Living-Harness, a self-evolving agent framework that dynamically updates its own execution harness to prevent recurring failures.
Researchers propose AgentSnare, a defense framework that uses deceptive environment observations to mislead and disrupt autonomous LLM pentesting agents.
A forensic audit of a radiology VLM benchmark found inconsistencies across datasets, DICOM rendering, prompts, APIs, and statistical code artifacts.
Research explores market designs for AI model training data, aiming to compensate human content creators beyond current 'free-for-all' or 'strong IP' models.
Euclid-MCP proposes a standardized protocol server for integrating LLMs with Prolog-based symbolic reasoning, aiming for reliable logical outputs.
Researchers propose a statistically-grounded method for sparse-feature intervention in LLMs, enhancing activation steering for behavioral control without fine-tuning.
Research highlights gaps in current LLM benchmarks, arguing they fail to measure analytical knowledge work and judgment critical for white-collar tasks.
Research introduces PUPPET, a taxonomy and resource to predict human belief change in dialogues with manipulative LLMs.
New research introduces an exact method for measuring state usage in selective state-space models (SSMs) like Mamba, detailing how information flows.
Research introduces TSCoNet, a two-stage Copula CNN-LSTM model for uncertainty-aware spatio-temporal forecasting of correlated environmental variables.
Research proposes a two-rate error measurement for LLM protocols to audit correction vs. corruption, improving understanding of their impact.
Research finds automated evaluation of LLM agents is unreliable, with errors propagating through tool-use chains. Benchmarked 9 LLMs.
InfiniteScienceGym is a new procedurally generated benchmark for evaluating LLMs on scientific reasoning from empirical data, aiming to overcome biases in human-curated datasets.
Research identifies decision boundary proximity as a common cause for miscalibrated confidence and paraphrase sensitivity in medical Vision-Language Models.
Amex Ventures invested in Fazeshift, an AI startup deploying autonomous agents to execute end-to-end accounts receivable workflows.
OpenAI's Daybreak Red and Daybreak Blue cyber defense models are now on Amazon Bedrock with chip-level zero-operator data access.
Major tech firms are shifting from post-hoc AI detection to cryptographically embedded content provenance and traceability standards.
Databricks open-sourced Metals v2, a Java and Scala language server designed to help AI agents navigate multi-million line codebases.
Anthropic introduced output watermarking for Claude models to comply with the EU AI Act Code of Practice on Transparency.
Google Cloud integrates Looker's semantic layer with Gemini Enterprise to govern NL2SQL queries and prevent metric hallucinations.
AWS released a production reference architecture for self-hosting a governance gateway between Claude developer tools and Amazon Bedrock.
Anthropic canceled planned Claude price increases following OpenAI's recent API price cuts of up to 80%.
Security researchers used fewer than 20 AI prompts on public models to uncover a now-patched zero-day vulnerability in Zoom.
Allvue Systems launched Intelligent Loan Operations, using AI agents to automate loan transaction data ingestion into general ledgers.
NVIDIA released Nemotron 3.5 Lightning, a specialized model optimized for fast tool calls, result validation, and subagent execution.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion