[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???
Key Google DeepMind pioneers including Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are departing or transitioning roles.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Key Google DeepMind pioneers including Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are departing or transitioning roles.
Researchers proposed an adaptive training controller for Conditional Value-at-Risk (CVaR) Risk-Aware Q-Learning to stabilize finite-budget training.
Researchers propose 'oblivious audits' to prevent model providers from detecting and manipulating regulatory or compliance evaluations.
Researchers identify a critical ML failure mode where models reconstruct proxy equations instead of learning robust underlying features.
Researchers propose Elbow-Based MoE Routing, a training-free inference plugin that dynamically selects Mixture-of-Experts active counts.
An academic paper challenges the reliability of ranking anomaly detection algorithms, citing non-aligned benchmark settings across research.
Researchers propose EvolveNet, a framework for LLM agents to self-improve by evolving their execution harnesses rather than model weights.
A research paper benchmarks LLMs fine-tuned via QLoRA, finding that superior financial sentiment scoring does not guarantee trading returns.
Researchers propose PriDyG, a framework combining GNNs and LLMs for edge-level differentially private inference on dynamic graphs.
Researchers introduce Trident, demonstrating that autonomous Deep Reinforcement Learning cyber defenses are highly vulnerable to adaptive agents.
Researchers find task-vector subtraction in vision-language-action models causes global performance collapse rather than targeted skill removal.
An academic study demonstrates that choices in LLM inference frameworks introduce behavioral and benchmark variability for identical models.
Researchers propose a framework using economic decision theory axioms for label-free evaluation and regularization of LLM reasoning.
Researchers propose a data-aware, scalable sensitivity analysis method to detect unfair feature influence in decision tree ensembles.
Researchers demonstrated a method using CLIP to execute universal, targeted adversarial attacks on models without needing training data.
Researchers propose using cryptographic fuzzy extractors to prevent model inversion attacks on facial biometrics and embedding vectors.
An evaluation of code language models shows security patch detection is heavily reliant on commit messages rather than code analysis.
Research reveals multilingual LLM evaluation gaps on MGSM are heavily distorted by token output caps rather than actual reasoning limits.
Researchers introduced a framework using semantically equivalent adversarial attacks to expose intrinsic hallucinations in RAG systems.
Researchers introduced FinReportBench, an expert-grounded benchmark using a 35-item rubric to evaluate institutional financial reports.
Researchers identified 'referential dangling,' a paradigm-level failure mode where prompt compression splits and deletes dependent text pairs.
Researchers introduce MirageBench, showing that personalized LLMs with persistent memory fabricate user profiles beyond evidence.
Researchers propose EASy, an agentic orchestration framework designed to optimize execution efficiency and compute cost alongside task success.
A new research benchmark, Skill-Use, evaluates LLM agents on their ability to independently recognize and apply structured execution skills.
Research reveals LLM confidence estimation techniques are highly sparse, with models like Qwen3-32B clustering most outputs at exactly 95%.
Research addresses the trade-off in latent chain-of-thought models where continuous states lower inference costs but eliminate readable reasoning traces.
Researchers evaluate verbalized confidence in small language models (0.5B-14B) to determine if they can safely defer to human reviewers.
A research paper demonstrates that language models fail to adapt their reasoning when evaluated on varying modal logic constraints.
Researchers introduce Skill Entropy, a new metric to measure how LLMs transition between distinct skills in multi-step reasoning tasks.
Researchers introduced FinProBench, an agentic evaluation framework using rubrics derived from professional practitioner deliverables.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion