A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
A research paper proposes a knowledge-injection framework to improve zero-shot delirium prediction using LLMs and electronic health records.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A research paper proposes a knowledge-injection framework to improve zero-shot delirium prediction using LLMs and electronic health records.
LeanFlow is an LLM agent system designed to translate mathematical papers into formal, buildable Lean projects, evaluating runtime mechanisms.
A research paper introduces STeMP, a spatio-temporal modeling protocol to address data sensitivity and methodological choices in ML model quality estimation.
Research paper introduces Fisher widths to measure local parameter fluctuations and provide anisotropic recovery in statistical manifolds.
MiniCache is a reusable program caching framework that converts Program-of-Thought (PoT) programs into parameterized cache objects for efficient LLM inference.
Research proposes Minimum Bayes Risk decoding for Error Span Detection in automatic machine translation evaluation, improving error localization.
Research proposes sparse autoencoders (SAEs) to create interpretable embeddings for analyzing large text corpora, addressing limitations of LLM-based and dense embedding methods.
Research explores compiling discrete programming structures, like conditionals and iteration, into differentiable recurrent neural networks.
A survey of deepfake generation and detection techniques identifies various deepfake types, taxonomies of methods, datasets, and top detectors.
Research finds open-weight instruction-tuned LLMs consistently reproduce homogeneity bias across models, decoding settings, and identity signals.
Research proposes a method for comparing generative models using KL divergence to quantify uncertainty in model evaluation, improving rigor.
Research introduces Pairwise Quantile Regression, extending conditional distribution modeling for high-dispersion variables beyond standard least squares.
Research explores using Quantum Kitchen Sinks (QKS) for RF spectrogram anomaly detection to secure wireless spectrum management.
Research introduces Exo-MDPs, a structured class of Markov Decision Processes for sample-efficient reinforcement learning in operations research.
Research introduces heat-kernel entropy profiles for weighted measures on manifolds to account for geometric support in effective sample size.
Research presents enhanced Membership Inference Attacks (MIAs) on diffusion models using a frequency-domain approach, improving data privacy breach detection.
Research suggests long-horizon time-series forecasting benchmarks may not require complex transformer architectures, finding simpler linear models competitive.
Research characterizes the VC dimension and sample complexity of Transformers, providing tight bounds for depth-L models with W parameters and sequence length T.
Researchers introduced ER-JEPA, a hierarchical self-supervised learning framework for multivariate time series, demonstrated in ECG analysis.
ConfidenceBench evaluates 15 frontier LLMs on verbalized confidence calibration using the Brier score, addressing risks of fluent but incorrect answers.
Research proposes a category theory approach using Kan extensions to define and compute structural invariants for transfer learning between tasks.
Research introduces VPWEM, a visuomotor policy using working and episodic memory to handle non-Markovian robotic tasks, addressing long-term memory limitations.
Research investigates when weight-tied looped transformers implement algorithms, finding a linear computation frontier dictated by training budget.
Research explores parameter-free RDB encoders for foundation models to predict missing values across diverse enterprise tasks without retraining.
Research finds error amplification limits converting Artificial Neural Networks (ANNs) to Spiking Neural Networks (SNNs) for continuous control.
Research demonstrates scaling closed-loop LLM-based channel configuration for neural networks by optimizing widths via executable code generation.
Research introduces CT-Merging, a method to combine multiple LoRA adapters into one multi-task adapter, reducing storage and inference complexity.
Research proposes a shift from reactive, error-patching AI maintenance to proactive, test-driven development for deployed systems to improve generalization.
StabilityBench is a new research benchmark evaluating LLM instability and context dependence, particularly in high-stakes environments.
Chronofy, a new RAG architecture, introduces a temporal-logical decay mechanism to prevent LLMs from using obsolete facts, addressing 'temporal hallucination'.
Research explores reliable extraction of "strong lottery tickets"—sparse, performant subnetworks in large neural networks prior to training.
Research identifies 'attention specific heat' as a reliable predictor for Grokking, improving neural network generalization and training efficiency.
SevDiff is a new diffusion model for generating rare, high-severity conflict trajectories in ADAS evaluation, addressing bias in real-world data.
A new deep learning framework, JuGAAD, uses geospatial and census data to downscale socioeconomic indicators in India, addressing data resolution mismatches.
Research proposes using neural predicates to formalize and scale view generation in the Black-Litterman model for portfolio construction, reducing subjectivity.
Research proposes a statistical framework for optimal noise-level allocation in diffusion model training, moving beyond heuristic schedules.
Research uses multi-agent reinforcement learning (MARL) for decentralized conflict resolution in air corridors under degraded surveillance.
A new research paper introduces Monkey King Bang, a multimodal foundation model designed for unified scientific reasoning across diverse domains and data types.
Research explores methods to accelerate the training of Masked Diffusion Language Models (MDMs), which are slower to learn than autoregressive models (ARMs).
ReliableTableQA is a framework training LLMs to annotate the statistical reliability of tabular QA results, assessing if answers are statistically meaningful.
Research paper retracts prior claims of improved long-context retrieval for LLMs using Surprisal-Aware Residual Test-Time Training (SR-TTT), citing evaluation artifacts.
Research details MuonQ, an optimization method to improve low-bit quantization for the Muon optimizer in large language models by preserving directional fidelity.
PISmith, a reinforcement learning framework, is proposed to red team and systematically assess prompt injection defenses for LLM applications.
Research describes a 'self-evolving' recommendation system using LLM agents for autonomous model optimization, reducing manual iteration.
Variational Speculative Decoding (VSD) is a new method for training draft models in speculative decoding to maximize target-model acceptance.
Research proposes reusing internal audio language model representations for temporal localization to improve speed and accuracy over token generation.
Researchers propose Multimodal CoLRAG-TF, a quadruple-fusion RAG architecture combining text, keywords, knowledge graphs, and image similarity for complex PDFs.
NeuraLSP proposes a neural spectral preconditioner to accelerate solving sparse linear systems from PDEs, improving on traditional multigrid methods.
Physics-Informed Learning via Diffusion (PILD) is a new framework that integrates physical constraints into generative diffusion models using a probabilistic residual formulation.
Researchers propose Amortized Group Relative Policy Optimization, a reinforcement learning technique for Diffusion Large Language Models (dLLMs).
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion