How evals drive the next chapter in AI for businesses
OpenAI publishes guidance on using evaluations (evals) to measure and improve AI performance in business deployments.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
OpenAI publishes guidance on using evaluations (evals) to measure and improve AI performance in business deployments.
OpenAI and Target partner to launch a Target ChatGPT app for personalized shopping and expand ChatGPT Enterprise internally.
OpenAI published a system card for GPT-5.1-CodexMax, detailing model-level safety training and product-level mitigations like sandboxing.
OpenAI launches GPT-5.1-Codex-Max, a faster agentic coding model optimised for long-running, project-scale software tasks.
Scania claims productivity gains and accelerated innovation by deploying ChatGPT Enterprise with guardrails across its global workforce.
Google DeepMind announced new Gemini 1.5 Pro features, including an updated context window and native audio understanding, through a new API.
Google DeepMind establishes a new research lab in Singapore, focusing on AI advancement in the Asia-Pacific region.
The rapid advancement from GPT-3 (2020) to Gemini 3 (anticipated) highlights accelerated AI capabilities, moving from chatbots to agents.
Google DeepMind announced Gemini 3, a new generation of multimodal AI models, with limited details on capabilities or release timelines.
Latest research covers hardware-aware quantization for model efficiency, model lineage tracing for governance, and task-oriented grasping in robotics.
Intuit and OpenAI formed a multi-year partnership exceeding $100M for Intuit app integration into ChatGPT and broader use of OpenAI models.
Google DeepMind released WeatherNext 2, an AI model claiming more efficient, accurate, and higher-resolution global weather predictions.
OakNorth Bank launched a dedicated informational page designed specifically for AI crawlers to retrieve verified corporate data.
Hugging Face announced easier building and sharing of ROCm kernels, potentially improving AMD GPU integration for AI workloads.
Google's Android team reports memory safety vulnerabilities fell below 20% in 2025 by adopting Rust for new code.
Google DeepMind's SIMA 2 is a Gemini-powered AI agent designed to play, reason, and learn within virtual 3D environments.
OpenAI claims a new sparse model approach improves mechanistic interpretability of neural networks, enhancing transparency and reliability.
State of AI's latest research compilation covers efficient long sequence decoding, multimodal video generation, and neuro-symbolic CoT validation.
OpenAI releases GPT-5.1 via API with faster reasoning, extended prompt caching, better coding, and new shell/patch tools.
Philips deployed ChatGPT Enterprise to train 70,000 employees in AI literacy and responsible use across healthcare operations.
OpenAI opposes NYT subpoena seeking 20M user ChatGPT conversations, citing privacy; accelerating data protection measures.
The concept of 'AI job interviews' evaluates AI model performance through simulated role-based tasks, beyond standard benchmarks.
OpenAI releases GPT-5.1, a GPT-5 series upgrade with improved conversational tone and user-facing customization options.
OpenAI published a system card addendum for GPT-5.1 Instant and Thinking, covering updated safety evals including mental health and emotional reliance.
Google DeepMind research details how AI visual perception differs from human perception, impacting object recognition and scene understanding.
Anthropic secured $13 billion in funding, with commentary suggesting the investment emphasizes AI safety and potential future regulation.
Google DeepMind pilot in Northern Ireland schools with Gemini and other generative AI tools saved teachers 10 hours weekly.
OpenAI claims new guardrails for ChatGPT reduce misinformation and harmful content risks, improving trust in the platform.
Research on deformable object dynamics, efficient inference, and reliable simulation indicates advances in modeling complex physical interactions for robotics and AI.
Anthropic's reported payout in legal dispute highlights growing pressure on AI developers regarding creator rights and copyright. Broader implications for model training data use.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion