Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic Uncertainties
Researchers propose combining mean-variance estimation with Bayesian neural networks to separate data noise from model uncertainty.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose combining mean-variance estimation with Bayesian neural networks to separate data noise from model uncertainty.
Researchers prove modern neural networks can be factored into a nonlinear singular value decomposition without altering input-output behavior.
Research shows model quantization silently breaks algorithmic recourse and counterfactual explanations in decision-making systems.
Researchers propose MARGIN, an online calibration method to normalize and correct incomparable confidence scores across multi-model agent pools.
Researchers propose APPO, an RL method for agentic LLMs that improves credit assignment over intermediate tool-calling decisions.
Researchers propose ECHO, a selective memory pruning and tracing framework to optimize context windows for long-horizon agentic RL.
Researchers propose a Minimum Description Length framework for layer-adaptive optimization to prune and allocate LLM capacity efficiently.
Researchers formalized strategic auditee gaming under continuous compliance monitoring, identifying patterns of delay and report drifting.
Researchers formalized a game-theoretic model showing how developers can strategically manipulate AI systems to evade privacy-preserving audits.
Researchers introduce Chain-of-Models, a pipeline where a secondary LLM audits the reasoning of LLM judges to reduce cognitive bias.
Research shows knowledge distillation from larger models like Gemma-2-9B to smaller models improves context accuracy but worsens bias calibration.
Researchers introduce TokenSwap to benchmark and mitigate the 'modality gap'—inconsistent outputs in multimodal LLMs for equivalent inputs.
Research identifies systematic failures in LLM-as-a-judge evaluators, which mistake structural formatting for factual truth under adversarial load.
Researchers find that traditional evaluation metrics like perplexity are unreliable for federated pre-training, complicating downstream performance prediction.
An academic paper introduces a benchmark testing LLM structural reasoning and multi-step logic over complex, long-horizon financial statements.
Researchers introduced Self-Supervised Skill Optimization (SSO), a framework for LLM agents to learn reusable skills from unlabeled data.
Researchers introduced a meta-evaluation framework that audits LLM benchmark datasets at the sample level across five latent dimensions.
An arXiv research paper argues that financial LLM validation must shift from model-centric benchmarks to system-level evaluations.
Researchers propose TextCloak, using reinforcement learning to inject perturbations into text to prevent unauthorized LLM training.
Researchers propose an attribution-guided steering method to diagnose and mitigate sycophancy in LLMs at the individual token level.
Researchers propose 'Mixture-of-Translators' to translate KV caches across heterogeneous LLMs, enabling context reuse between model architectures.
Researchers identify universal 'futile reasoning' in LLMs, where models generate expensive, incorrect chains of thought on hard tasks.
Researchers propose InMyStyle, a system using local helper LLMs to fine-tune small, per-user LoRA adapters to match individual writing styles.
Researchers introduced Data Turnstile, an open-source framework generating high-quality function-calling data to fine-tune small models.
Researchers introduce CalibratedRubric, a task-adaptive LLM evaluation framework using Bayesian filtering and item response theory.
Researchers introduce Zero-Mem, a method enabling LLM agents to perform memory operations without generating intermediate tokens.
Researchers introduced PTP, a method for reconstructing LLM system prompts from model outputs using previous-token prediction.
Academic research identifies a "knowledge utilization" failure where LLMs fail to act on user preferences despite having them in context.
Academic study demonstrates that pretraining LLMs on interventional data fails to correct causal reasoning errors caused by observational bias.
Academic study proves vision-language models exhibit sycophancy in cooperative tasks, failing to correct user errors.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion