Adversarial observations in probabilistic State-Space Models for robust Reinforcement Learning
Researchers analyzed how adversarial observation attacks bypass detection by remaining statistically consistent with linear state-space models.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers analyzed how adversarial observation attacks bypass detection by remaining statistically consistent with linear state-space models.
A study of the MOOVE platform reveals that blinded pairwise preferences of LLM outputs fail to serve as reliable indicators of clinical safety.
Researchers introduce JudgeArena, a unified, reproducible framework to evaluate and standardize LLM-as-a-judge methodologies.
A new study reveals LLMs frequently generate correct legal answers while citing incorrect or hallucinated statutory authority.
Researchers release BODHI, investigating whether RLVR-trained LLMs actually expand reasoning boundaries or merely improve sampling efficiency.
Researchers identify a 'comprehension-containment decoupling' where LLM safety alignment fails on low-resource language slurs.
Researchers introduce the LLM Nominal Response Model (LLM-NRM), a psychometric framework analyzing incorrect MCQ choices to evaluate model behavior.
Researchers introduce TQLite, utilizing multi-LLM jury-guided distillation to train smaller models for real-time translation evaluation.
Researchers propose a multidimensional framework evaluating LLM statistical reasoning through accuracy, explanation structure, and lexical similarity.
Researchers introduce Iterative Context Optimization (ICO), a semantic-shift jailbreak technique that bypasses safety alignments.
Researchers introduced a signed-network framework that injects explicit relational priors into multi-agent systems to guide agent convergence.
Academic study evaluating 23 LLMs reveals that standard commonsense benchmarks fail to reliably predict performance on downstream tasks.
An academic study finds LLMs fail at active multi-turn information gathering and knowing when to stop querying during reasoning tasks.
Researchers propose DUD, a new method for uncertainty quantification in LLMs by decoupling attention and feed-forward hidden state updates.
Researchers evaluate cross-lingual bias in GPT-5.2 and Gemini 2.5 Flash using 4,900 English-Swahili prompt pairs to test alignment gaps.
Researchers demonstrate byte-level language models enable exact knowledge transfer because they share a uniform output space independent of tokenization.
Researchers introduced an efficient knowledge distillation method using offline top-K logits and a fused chunked KL loss to train small LLMs.
Researchers introduced M-GATE, a multilingual benchmark evaluating grammatical proficiency, translation accuracy, and efficiency across 30 languages.
Researchers introduced VIBE, a benchmark using Valence-Arousal-Dominance to measure subtle emotional and political bias in LLM outputs.
Researchers find that model sensitivity to noise, causality of predictions, and where repairs can occur dissociate across layers in LLMs.
A new academic study, SciRet, evaluates the compute-to-performance tradeoffs of multi-stage RAG pipelines across scaling document corpora.
Researchers introduced PAST-Bench, a benchmark evaluating how personal AI agents leverage history and skills to self-improve over time.
Researchers introduced 'WorldCup Arena' to evaluate frontier LLMs on live forecasting tasks to eliminate data leakage and memorization.
Researchers propose AnchorKV, a compression scheme claiming to reduce LLM key-value cache memory footprints by 20x without discarding tokens.
Researchers introduced DP-MemView, a differentially private memory interface designed to prevent LLM agents from leaking sensitive user attributes.
An arXiv paper exposes how LLM benchmark improvements often reflect changes in generation search paths rather than actual capability gains.
Researchers propose a self-distilled reward shaping method to solve the credit assignment problem in agentic reinforcement learning.
Researchers propose a method to detect LLM reasoning failures by analyzing the dynamic evolution of Chain-of-Thought traces.
Researchers introduce a distractor-aware truncation protocol to isolate how context length and noise affect LLM benchmark performance.
Researchers propose a search-rubrics reranking method to select document sets based on diversity and authority, rather than individual relevance.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion