Active Inference as a Convex Markov Decision Process
Research frames Active Inference (AIF) as a convex Markov Decision Process, simplifying policy optimization for expected free energy minimization.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research frames Active Inference (AIF) as a convex Markov Decision Process, simplifying policy optimization for expected free energy minimization.
Research finds optimizing reward model inference speed in RLHF pipelines significantly reduces bottlenecks, suggesting C++ implementations over PyTorch.
Researchers propose Differentially Private Decoupled Training (DP-DT) to improve utility in differentially private neural network training by decoupling representation learning from privacy enforcement.
OPIUM is a training-free method introduced to mitigate unintended externalities like over-refusal and weakened safety from LLM activation steering.
SCPP is a new open-source Python library that unifies the interface for various soft clustering methods, including fuzzy, probabilistic, and deep learning.
New research proposes an algorithm for reinforcement learning that achieves asymptotically optimal regret without dependence on the task horizon.
Research explored Time Series Foundation Models (TSFMs) like TimesFM, Chronos, and MOIRAI for zero-shot heart rate variability forecasting from consumer wearables.
Research explores Gaussian averaging as a smooth surrogate for quantized neural networks, deriving bounds on local oscillation for discontinuous models.
Koopman Dreamer proposes a spectrally constrained latent dynamics core for world models to improve stability and control in long imagined trajectories.
Researchers propose using physical noise as a native regularizer in photonic hybrid quantum neural networks, testing on Iris, Digits, and MNIST datasets.
Research explores expert-guided editing for time-series foundation model forecasts, allowing human feedback to revise model-generated trajectories.
New research on the Lattice Deduction Transformer (LDT) reveals neural 'solvers' make one-shot predictions rather than iterative deductions.
A Good Practice Guide for quantifying uncertainties in machine learning models applied to photoplethysmography (PPG) signals from wearables was published.
Research explores the frequentist consistency of Prior-Data Fitted Networks (PFNs) for causal inference, addressing uncertainty quantification.
Research evaluated the survival of quantum kernel geometry on IBM quantum hardware using a four-qubit ZZ feature-map kernel and air-quality data.
Research identifies a "Dominant-vs-Dominated" (DvD) imbalance in text-to-image diffusion models where one concept dominates multi-concept generation.
Researchers propose a neuromorphic-inspired Receptron model for efficient, non-linear classification at the edge, reducing compute and memory.
Research proposes a novel clustering method for LLM inference at scale, ensuring per-sample quality control and reducing cost and latency bottlenecks.
Research provides uniform confidence bands for Kernel Ridge Regression (KRR), improving inferential theory for nonstandard data applications.
Researchers introduced Directional Kernel Mean Difference (DKMD), a signed statistic for univariate distribution comparison, preserving directional shifts.
Researchers propose a novel differentially private training framework for deep neural networks that selectively privatizes inputs, not labels.
New research proposes Memory Merge DQN, an alternative target network update for Deep Q-networks to improve stability and preserve value function structure.
Research presents a scalable method using Determinantal Point Processes for diverse, high-quality data selection for large model training.
Research proposes a generative AI-enhanced probabilistic multi-fidelity surrogate modeling framework using transfer learning to address data scarcity.
Research introduces 'in-span learning' for reduced-order models (ROMs), allowing them to self-adapt and maintain accuracy when dynamics drift.
Research introduces Experience Augmented Policy Optimization (EAPO) to improve LLM reasoning via Reinforcement Learning with Verifiable Rewards (RLVR).
Research proposes aggregating Bayesian update rules as experts to achieve uncertainty-aware prediction on data streams with adaptive inferential choices.
Research finds DreamerV3-family model-based RL agents catastrophically forget tasks, with the 'actor' component, not the 'world model', being the source of forgetting.
Research proposes ECRAM, a continual learning method for edge devices, combining previous data summaries with new sensory input to avoid catastrophic forgetting.
Research proposes LAARA, a framework for adaptive rank allocation in parameter-efficient fine-tuning (PEFT), addressing the suboptimality of uniform rank.
Research demonstrates that intermediate layers of LLMs and speech models predict brain responses to language, identifying abstraction as key to alignment.
New research proposes an "in-the-flow" agentic system optimization for LLM planning and tool use, addressing scalability and generalization limits of monolithic policies.
A new benchmark, SHOVIR, is proposed to evaluate Vision-Language Models in radiology report generation, focusing on whether diagnostic statements are grounded in image evidence, not just lexical overlap.
Research introduces KoRe, a method to represent LLM knowledge externally, addressing opacity, debugging difficulties, and hallucination issues in parametric knowledge.
Research identifies 'self-preference bias' in rubric-based LLM evaluation where models favor outputs from themselves or their family, skewing benchmarks.
Research explores PortLLM, a training-free, data-free method for adapting LLMs, highlighting its short-term temporal portability for LoRA patches.
Researchers propose a meta-learning framework for aligning multilingual large language models by leveraging preference data across languages.
Researchers introduced 'vocabulary dropout' to prevent collapse in LLM co-evolution, enhancing curriculum diversity in self-play systems.
Researchers propose a graph-based RAG method to reduce hallucinations in complex question answering by improving context retrieval for LLMs.
Research investigates how human-AI co-authorship and large language model (LLM) edits impact the detection of an author's native language (L1) traces in non-native writing.
Research challenges the common practice of using emotion embedding similarity (e.g., emotion2vec) to evaluate emotional expressiveness in speech generation models.
Researchers propose BITEMBED, a low-bit framework for LLM-based text embeddings to improve efficiency and reduce storage and bandwidth overhead.
New research argues that current methods for evaluating LLM activation explanations are structurally flawed, failing to penalize individual false claims.
New research benchmark, CEO-Bench, evaluates AI agents on long-horizon, real-world tasks requiring complex skill orchestration and adaptation.
Researchers propose LaSEr-Edit, a method for localized, span-level editing of LLM outputs to enforce safety and consistency constraints more reliably than brittle instruction-based control.
Research explores prompt programming for LLM cultural bias and alignment in strategic decision-making and document engineering tasks.
Research finds self-supervision, not clinical supervision, primarily drives representational convergence in medical foundation models, affecting their clinical usability.
Researchers propose JailMeter, a new evidence-based framework to consistently evaluate jailbreak attack effectiveness on large language models.
Research identifies a benchmark leakage issue where 'out-of-distribution' classes were inadvertently included in training, skewing OOD detection metrics.
New research, HyGRL, proposes a hybrid graph reasoning framework for multi-entity questions, addressing limitations in conventional RAG and LLM-constructed Graph-RAG.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion