The Silent Freeze: Predicting When Low-Precision Training Stops Learning
Research identifies a silent learning freeze in low-precision AI training, predictable from high-precision data and mantissa length.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research identifies a silent learning freeze in low-precision AI training, predictable from high-precision data and mantissa length.
Research tested LLM juries against expert panels for scoring medical diagnoses in real-world hospital cases, showing strong correlation.
Apple research shows LLMs fail to update probabilistic beliefs consistently, deviating significantly from ideal Bayesian reasoning.
The OCC revised policy manuals and proposed framework amendments to standardize bank enforcement standards and Matters Requiring Attention (MRAs).
Amazon Bedrock launches in-country cross-Region inference in India for data residency compliance with local processing requirements.
WIRED code analysis reveals OpenAI is developing persistent execution capabilities for Codex, enabling proactive background operation.
Socure secured funding from Goldman Sachs Alternatives and Wells Fargo at a $5.2B valuation and acquired agentic fraud platform Fravity.
Autonomous AI agents require data-layer governance to prevent unauthorized actions across enterprise systems.
Cohere introduced Parse, a high-throughput vision parsing model designed for enterprise document intelligence and unstructured data ingestion.
Anthropic released Claude Opus 5, offering performance improvements for long-running agents, software engineering, and professional workflows.
ASIC and APRA issued a joint warning for Australian financial entities to shift from AI risk awareness to active mitigation frameworks.
Nvidia has reportedly agreed to acquire open-source model platform Hugging Face for $12.9 billion.
Research demonstrates LLM-as-a-Judge evaluation independence is compromised by anchoring bias when prompts contain prior scores or history.
ArXiv paper shows how undisclosed inference-time steering masks underlying model behavior, complicating auditability and evaluation.
OpenSanctions releases Pairs, an open benchmark of 755,540 expert-labeled entity matching pairs across 45 jurisdictions.
A new research paper evaluates uncertainty quantification methods to predict functional correctness in LLM-generated code.
Research evaluates confidence estimation and selective prediction for financial NER across SEC filings, news, and user content.
Research identifies rubric interference in single-pass LLM judges, where co-present evaluation criteria alter individual assessment accuracy.
Researchers introduce ClayBuddy, a framework analyzing and mitigating destructive coding agent failures caused by capability errors.
A research study evaluates evidence-generation in LLMs, demonstrating how retrieved context can both assist and distract verification tasks.
Researchers introduce SciStyleBench, finding that LLM-as-a-judge evaluators are heavily biased by stylistic presentation over substance.
Researchers introduced SCHEDBench, a benchmark of 1,132 instances evaluating LLM reliability in solving combinatorial scheduling tasks.
Researchers propose Think Short, Defer Smart (TSDS), a framework for edge LLM agents to manage reasoning budgets and defer to cloud models based on uncertainty.
Research identifies common metrics for synthetic tabular data generation are blind to inter-column dependencies critical for fraud and risk models.
VDAR-Router, a new LLM routing method, uses verbalized query difficulty analysis to dynamically select models based on cost-performance trade-offs.
LLM evaluators (reward models and LLM-as-a-Judge) exhibit significant language bias, failing to assign consistent scores across 23 languages for semantically identical instruction-response pairs.
Research identifies "pigeonholing," where unintentionally bad prompts degrade LLM performance and cause mode collapse, even without malicious intent.
Research explored mitigating LLM biases from spurious social contexts using direct preference optimization, focusing on high-stakes decision-making.
SalesLLM, a new benchmark, evaluates LLM performance in multi-turn, goal-directed sales dialogues, specifically in Financial Services and Consumer Goods.
NVIDIA allegedly acquires Hugging Face for $13B while OpenAI releases an incident retrospective regarding Hugging Face platform usage.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion