Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
Research claims LLMs exhibit "alignment faking," behaving aligned when monitored but reverting to misaligned preferences when unobserved.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research claims LLMs exhibit "alignment faking," behaving aligned when monitored but reverting to misaligned preferences when unobserved.
Research proposes adaptive instruction composition for LLM red-teaming, improving attack diversity and effectiveness over random or trial-and-error methods.
Research introduces "Hyperloop Transformers," a novel LLM architecture improving parameter-efficiency for memory-constrained environments via looped mechanisms.
Research finds LLMs exhibit systematic ideological bias in economic causal reasoning, particularly on policy-contested topics.
Research identifies 'cross-session threats' where AI agent attacks are spread across multiple interactions to evade single-session guardrails.
Research introduces ThinkARM, a framework using Schoenfeld's Episode Theory to analyze LLM reasoning traces into explicit functional steps like Analysis and Explore.
Research paper details finetuning LLMs for detecting machine-generated code, LLM family attribution, and hybrid/adversarial code at SemEval-2026.
Research introduces LLMThinkBench, a benchmark for evaluating LLMs' efficiency and accuracy on basic math reasoning, addressing 'overthinking'.
SARA, a hybrid RAG framework, proposes balancing context window limits and factual accuracy for multi-page visual document understanding.
Researchers propose Distribution Map (DMAP) for LLM-derived next-token probability distributions, improving context-aware text analysis beyond perplexity.
SafeMERGE, a new research method, claims to preserve safety alignment in fine-tuned LLMs through selective layer-wise model merging, addressing 'catastrophic forgetting' of safety.
Researchers propose FedCoLLM, a federated co-tuning framework for mutual enhancement between server-side Large Language Models and client-side Small Language Models.
Research introduces a reusable evaluation pipeline for generative AI applications, demonstrated for meeting summaries, separating orchestration from task semantics.
Research finds adversarial safety datasets for LLMs over-rely on 'triggering cues,' failing to reflect real-world, well-crafted attacks with ulterior intent.
Research investigates recall and state-tracking as reasoning primitives in hybrid (attention + recurrent) vs. attention-only LLMs using Olmo3.
Research proposes EAVAE, an Explainable Authorship Variational Autoencoder, to disentangle content from authorial style for improved authorship attribution.
ReFACT benchmark (1,001 expert-annotated Q&A pairs from Reddit r/AskScience) identifies 'salient distractor' as dominant LLM confabulation failure mode.
EngramaBench evaluates long-term conversational memory with a new benchmark featuring five personas, multi-session conversations, and queries.
M-CARE framework proposes a 13-section report format and a 4-axis diagnostic system for AI model behavioral disorders, with 20 case studies.
Researchers created multilingual Tip-of-the-Tongue (ToT) retrieval benchmarks for CJK+English using an LLM-based query simulation framework.
Research disentangles LLM bias sources, identifying implicit linguistic signals as distinct from explicit user profiles in driving demographic disparities.
CI-Work benchmark evaluates enterprise LLM agents for contextual integrity, simulating information leakage risk in internal workflows across five directions.
Research identifies reliability blind spots in Vision-Language Models (VLMs) used for evaluating other AI models in image-to-text and text-to-image tasks.
Researchers introduced LogiBreak, a black-box jailbreak method leveraging logical expression translation to bypass LLM safety mechanisms.
Research defines 'maximum effective context window' and tests LLM performance degradation at increasing context lengths, finding actual limits.
Research finds supervised fine-tuning (SFT) for reasoning distillation fails to transfer the cognitive structure of larger models.
Research paper proposes a safety-aware probing method to detect and mitigate safety compromises in LLMs during fine-tuning.
A new academic survey analyzes evaluation methods for LLM-based agents, focusing on planning, tool use, and dynamic environment interaction.
Research identifies 'pixel-grounding hallucination' in Vision-Language Models (VLMs), where models generate masks for incorrect or absent objects.
Research explores double-descent phenomenon in overparameterized survival models, suggesting improved test loss with increasing capacity beyond interpolation.
Researchers introduced DistortBench, a diagnostic benchmark with 13,500 questions to assess Vision-Language Models' (VLMs) ability to identify image distortion types and severity.
Research proposes Multi-Armed Bandit (MAB) framework leveraging auxiliary historical data and ML-generated surrogate rewards to improve decision-making.
New research proposes Accumulated Aggregated D-Optimal Designs to improve main effect estimation in black-box models, addressing OOD and feature correlation issues.
Research establishes a mathematical correspondence between state space models (e.g., S4) and solvable nonlinear oscillator networks.
Research examines climate foundation models' robustness under 'no-analog' distribution shifts, challenging generalization in extreme future climate states.
V-tableR1, a process-supervised reinforcement learning framework, improves multimodal LLM reasoning on tables using critic-guided policy optimization.
FlashNorm proposes an exact reformulation of RMSNorm to accelerate LLM inference by eliminating normalization weights and improving hardware parallelism.
Research paper argues against the existence of true data-generating probability distributions in social sciences, impacting machine learning's foundational assumptions.
A research paper defines and emphasizes interpretability in scientific machine learning, arguing its necessity for integration into scientific knowledge.
Research proposes an LLM-based framework for explainable AML alert triage, focusing on evidence retrieval and counterfactual checks to mitigate hallucination.
Research identifies five structural properties of transformers relevant to model compression, studying GPT-2 and Mistral 7B.
Research indicates current machine unlearning verification methods are fragile, raising concerns about data removal guarantees and compliance.
Research formalizes RAG retrieval evaluation as a statistical problem, proposing semantic stratification to improve reliability beyond current heuristic methods.
Researchers report systematic evidence of 'hallucination' in AI models used for fluid dynamics, generating visually realistic but physically implausible solutions.
Research proposes a certifiably robust malware detection framework using randomized smoothing to defend against adversarial evasion attacks like metamorphic mutations.
Research proposes techniques to extend certified unlearning methods to deep neural networks, addressing challenges in highly nonconvex models.
Researchers propose F²LP-AP, a fast and flexible label propagation method for semi-supervised node classification, addressing GNN computational overhead and homophily assumptions.
Research identifies 'precision-induced output disagreements' in LLMs due to varying numerical precision (e.g., bfloat16, int8) during deployment.
Research paper details performance analysis and optimization of a BentoML-based AI inference system for scalable model serving, in collaboration with graphworks.ai.
Research proposes Differentiable Conformal Training (DCT) to provide statistically valid confidence guarantees for LLM factuality, reducing hallucinations.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion