Evaluating chain-of-thought monitorability
OpenAI releases framework and 13-evaluation suite showing CoT reasoning monitoring outperforms output-only monitoring for AI control.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
OpenAI releases framework and 13-evaluation suite showing CoT reasoning monitoring outperforms output-only monitoring for AI control.
OpenAI and the U.S. Department of Energy (DOE) signed an MOU to collaborate on AI and advanced computing for scientific discovery.
OpenAI published a system card addendum for GPT-5.2-Codex, a coding-focused variant of GPT-5.2.
OpenAI announces GPT-5.2-Codex, a coding-focused model with long-horizon reasoning, large-scale code transformation, and cybersecurity features.
OpenAI releases GPT-5.2-Codex, a coding-specialized model with long-horizon reasoning, large-scale code transformation, and cybersecurity features.
Mistral AI introduced Mistral OCR 3, an optical character recognition model for document understanding.
Hugging Face and NVIDIA collaborate on NeMo Evaluator, an open evaluation standard for LLMs, benchmarking NVIDIA's Nemotron 3 Nano model.
Google DeepMind announced Gemini 3 Flash, a new frontier model optimized for speed and cost-efficiency with high intelligence.
OpenAI publishes data-driven report on enterprise AI adoption trends, tracking progression from experimentation to productivity gains.
NIST released draft guidelines focusing on mitigating cybersecurity risks when incorporating AI into organizational operations.
Google DeepMind released Gemma Scope 2, an open interpretability tool for the Gemma 3 model family, to aid AI safety research.
OpenAI launches FrontierScience benchmark to evaluate AI reasoning across physics, chemistry, and biology research tasks.
OpenAI introduces an evaluation framework for AI-accelerated biological research, using GPT-5 to optimise a molecular cloning protocol.
OpenAI launched new ChatGPT Images with improved image generation, faster performance, and precise editing, available in ChatGPT and API as GPT-Image-1.5.
Hugging Face released CUGA, an open-source framework for building configurable AI agents, aimed at democratizing agent development.
The 'State of AI' report highlights research in decentralized LLM serving, trustworthy decision support, and interpretable sparse autoencoders.
Google DeepMind announced improved Gemini audio models, enabling more powerful voice experiences and enhanced multimodal capabilities.
BBVA deploys ChatGPT Enterprise to all 120,000 employees in multi-year OpenAI partnership targeting AI-native banking.
OpenAI claimed their internal team developed Sora for Android in 28 days using Codex for AI-assisted coding and project workflows.
BNY deployed OpenAI-powered platform 'Eliza' enabling 20,000+ employees to build AI agents across the enterprise.
llama.cpp adds experimental model management functionality for dynamically loading and unloading models, improving resource efficiency.
OpenAI claims GPT-5.2 sets new benchmarks on GPQA Diamond and FrontierMath, including solving an open theoretical problem.
Google DeepMind and UK AI Safety Institute (AISI) deepen collaboration on AI safety and security research, focusing on critical infrastructure and national security.
Codex is open-sourcing AI models, as announced on the Hugging Face blog.
Disney licenses 200+ characters to OpenAI's Sora for fan videos; Disney also adopts ChatGPT Enterprise and OpenAI API company-wide.
OpenAI announced GPT-5.2, a new model in the GPT-5 series, confirming consistent safety mitigations and data sources.
OpenAI announces GPT-5.2, claiming improved reasoning, long-context, coding, and vision for agentic workflows via ChatGPT and API.
The Chrome Root Program and CA/Browser Forum are phasing out 11 legacy domain validation methods for HTTPS certificates, increasing web security.
Google DeepMind announced deeper collaboration with the UK government on AI safety, security, and prosperity initiatives.
OpenAI published an article outlining its approach to strengthening cyber resilience in advanced AI models, detailing risk assessment and misuse limitation.
Scout24 deployed a GPT-5-powered conversational search assistant for real-estate listings, using clarifying questions and tailored recommendations.
Mistral AI released Devstral 2, a new code generation model, and Mistral Vibe CLI for model interaction and fine-tuning.
Google DeepMind released FACTS, a benchmark suite to systematically evaluate large language models' factuality across multiple domains.
OpenAI co-founds the Agentic AI Foundation under the Linux Foundation and donates AGENTS.md to advance open standards for safe agentic AI.
METR Research details common elements of frontier AI safety policies from model developers, including evaluation, info security, and deployment safeguards.
OpenAI appointed Denise Dresser as Chief Revenue Officer to lead global revenue strategy across enterprise and customer success.
Commonwealth Bank of Australia deploys ChatGPT Enterprise to 50,000 employees via OpenAI partnership for customer service and fraud response.
OpenAI and Deutsche Telekom partner to deploy ChatGPT Enterprise for DT employees and multilingual AI products across Europe.
Google's Chrome security team details its approach to securing agentic capabilities, following the launch of Gemini in Chrome.
OpenAI and Instacart integrate grocery shopping and Instant Checkout into ChatGPT via deepened partnership.
OpenAI publishes internal enterprise data claiming accelerating AI adoption and productivity gains across industries in 2025.
Virgin Atlantic CFO describes using OpenAI tools to accelerate development, improve decisions, and enhance customer experience.
Basel Committee consults on making Pillar 3 disclosures machine-readable, standardizing bank data for automated analysis.
Hugging Face released swift-huggingface, a Swift client library for interacting with its platform, enabling Swift-native ML workflows.
OpenAI launched "OpenAI for Australia" to develop sovereign AI infrastructure, upskill 1.5M workers, and boost the country's AI ecosystem.
Research explored improving AI reasoning over long-tail data distributions and identifying critical points in complex AI systems for stability.
Hugging Face demonstrated using Claude to fine-tune an open-source LLM, combining proprietary model instruction with open-source flexibility.
Android expands a pilot program using Google AI to detect and protect users from in-call financial scams within financial apps.
OpenAI acquires Neptune, an experiment tracking and model monitoring platform, to enhance ML observability tooling.
OpenAI testing 'confessions' training method to make models self-report errors and undesirable behaviour.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion