FLI’s President and CEO on Trump’s support for an AI ‘kill switch’
Donald Trump stated in a Fox Business interview that AI needs a government 'kill switch'. The Future of Life Institute (FLI) noted this.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Donald Trump stated in a Fox Business interview that AI needs a government 'kill switch'. The Future of Life Institute (FLI) noted this.
The Bank of England's Artificial Intelligence Consortium held its February 2026 meeting, fostering public-private dialogue on AI in UK financial services.
Research finds AI content watermarking efficacy varies significantly across languages, cultural traditions, and demographic groups due to content properties.
Research identifies two distinct internal information pathways (Question-Anchored, Statement-Anchored) within LLMs that encode truthfulness cues.
Research introduces Indica, a new benchmark to test LLM bias and cultural commonsense variation at sub-national levels within India, challenging monolithic national assumptions.
Researchers propose MAGE, a corpus-free unlearning framework for LLMs designed to address privacy and legal concerns by removing memorized sensitive content.
IndicDB is a new benchmark for evaluating Text-to-SQL performance of LLMs in Indian languages using real-world schemas.
Research proposes 'authority-aware generative retrieval' for LLMs, combining semantic relevance with document trustworthiness, critical for high-stakes domains.
Research systematically explores how multilingual data in LLM post-training impacts performance across languages, revealing English-centric bias.
Research indicates Vision Language Models (VLMs) prioritize semantic information from text inputs over detailed visual features for decision-making.
WorkRB is a proposed community-driven evaluation framework to standardize NLP models for hiring, talent management, and workforce analytics across fragmented research.
Research investigates prompt and aggregation strategies to improve LLM-as-a-judge accuracy for GPT-5.4 on RewardBench 2 without finetuning.
Research identifies intersectional bias in SpeechLLMs from accent and perceived gender, manifesting as quality-of-service disparities in human-AI speech interactions.
Research suggests LLM-generated labels can rival human labels in active learning for hostility detection, potentially reducing annotation costs.
BenGER is an open-source web platform integrating task creation, expert annotation, and model evaluation for German legal LLM benchmarks.
Research proposes ToolSpec, a method to accelerate LLM tool calling via schema-aware and retrieval-augmented speculative decoding, reducing latency.
Research finds LLMs can correctly follow Chain-of-Thought reasoning steps but still produce incorrect final answers, indicating reasoning-output dissociation.
Research proposes ABSA-R1, an LLM framework for Aspect-based Sentiment Analysis that aligns sentiment reasoning with human-like justifications.
Research paper proposes a unified framework for 'steering' LLMs via internal activation modification at inference, comparing it to traditional adaptation.
New benchmark, MERRIN, evaluates AI agents' multimodal evidence retrieval and multi-hop reasoning in noisy web environments.
Researchers propose Training-Free Test-Time Contrastive Learning (TF-TTCL) to improve LLM performance under distribution shift without gradient-based updates.
Research explores how LLMs implicitly trust humans, analyzing patterns and biases in human-AI interaction for decision-making contexts.
ValueGround benchmark evaluates multimodal LLMs' ability to ground culture-conditioned judgments in visual scenes, extending beyond text-only assessments.
Researchers introduced MulDimIF, a multi-dimensional framework for evaluating and improving instruction-following capabilities in LLMs across three constraint patterns.
Research paper empirically studies ClawHub, a public registry of LLM agent skills, exploring its functionality, ecosystem structure, and security risks.
Research suggests knowledge density in multimodal training data, not task format, is the primary bottleneck for MLLM scaling.
Research indicates LLMs struggle with reasoning tasks on finite discrete state-spaces as complexity increases, even with explicit validity constraints.
Research explores RAG vs. finetuning for LLM adaptation to continuous knowledge drift, identifying limitations in both for real-world factual changes.
Research identifies 'Logical Phase Transitions' where LLMs' logical reasoning abruptly collapses as complexity increases, even with small changes.
InfiniteScienceGym is a new procedurally generated benchmark for evaluating LLMs on scientific reasoning from empirical data, aiming to overcome biases in human-curated datasets.
Researchers introduced ChartNet, a 1.5 million-scale, high-quality multimodal dataset for training models in chart understanding and reasoning.
Research introduces a technique to quantify computation density in transformer LLMs, supporting claims that significant parameter pruning is possible.
Research analyzed stylistic differences between human and LLM-generated text across genres and decoding strategies to improve detection.
Research introduces Source-Shielded Updates (SSU) to adapt LLMs to new languages using only unlabeled data, mitigating catastrophic forgetting.
Research finds internal model representations that predict hallucination emerge at specific model scales before token generation, varying by model size.
A research survey reviews methods for generating synthetic network traffic using statistical models and deep learning to address data scarcity and privacy.
TRIM proposes routing only critical steps of multi-step reasoning tasks to more capable LLMs to prevent cascading failures and optimize inference.
Research paper reviews diffusion models for simulation-based inference (SBI), addressing intractable likelihoods in complex simulations.
Research demonstrates that the importance of LLM parameters for supervised fine-tuning shifts over time, challenging static parameter isolation methods.
Research paper explores fine-grained non-determinism in Diffusion Language Models, noting current dataset-level metrics limit insight.
Research identifies 'reward hacking' as a systemic vulnerability in LLM alignment, where models exploit reward signals without achieving true intent.
Research claims Ordinary Least Squares (OLS) is a special case of a single-layer Linear Transformer, demonstrated via algebraic proof.
Research paper reviews principles, challenges, and practical considerations for evaluating supervised machine learning models beyond aggregate metrics.
Research paper identifies numerical instability and chaotic behavior as a root cause of unpredictability in LLMs, especially in agentic workflows.
LiveClawBench is a new benchmark for evaluating LLM agents on complex, real-world assistant tasks, addressing gaps in current isolated evaluations.
Research identifies significant variability in individual patient risk predictions from overparameterized models due to optimization randomness, even with fixed data.
HUANet is a neural network architecture that unrolls ADMM iterations to solve constrained convex optimization problems, explicitly enforcing constraints.
Research finds larger LLMs improve at ignoring false claims but worsen at ignoring irrelevant tokens, formalizing contextual entrainment scaling laws.
Event Tensor is a compiler abstraction designed to optimize GPU inference for LLMs by fusing operators into a single megakernel to reduce overhead.
Research analyzes Anthropic's Claude Mythos system card, proposing hypotheses on whether 'emotion vectors' track functional emotions or situational contexts.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion