PaperBench: Evaluating AI’s Ability to Replicate AI Research
OpenAI releases PaperBench, a benchmark measuring AI agents' capacity to replicate published AI research end-to-end.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
OpenAI releases PaperBench, a benchmark measuring AI agents' capacity to replicate published AI research end-to-end.
OpenAI submitted a response to the UK government's AI copyright consultation, advocating for pro-innovation policy positions.
OpenAI raises $40B at $300B valuation to scale compute infrastructure and advance AGI research.
Hugging Face detailed its approach to secrets management, integrating Vault, Kubernetes, and SOPS to secure credentials across AI infrastructure.
Hugging Face claims accelerated LLM inference performance using Text Generation Inference (TGI) on Intel Gaudi hardware.
Basel III monitoring report shows increased risk-based capital ratios and stable leverage and NSFR for large international banks.
Google DeepMind announced Gemini 2.5, claiming it is their most intelligent AI model with built-in 'thinking' capabilities.
OpenAI integrates native image generation into GPT-4o, enabling multimodal output within a single model API.
OpenAI published a system card addendum for GPT-4o's native image generation, noting photorealistic output and image-to-image transformation.
Hebbia claims its AI platform automates 90% of finance and legal work using OpenAI models for deep document research.
METR Research is developing a new benchmark to evaluate AI models' ability to complete complex, multi-step tasks requiring long-term memory and planning.
Hugging Face published its response to the White House RFI on AI, advocating for open-source AI and balanced regulation.
Eugene Yan and Chip Huyen will present best practices for building LLM-powered applications at NVIDIA GTC 2025.
NVIDIA announced new open models and datasets for physical AI developers at GTC 2025, expanding their robotics and simulation ecosystem.
OpenAI announced March 2025 ChatGPT for Business updates: enhanced agentic features, team customization, and interactive capabilities.
Mistral AI released 'Mistral Small 3.1', an updated compact model intended for production use cases, positioned as a highly efficient alternative.
METR Research suggests priorities for the White House Office of Science and Technology Policy's (OSTP) AI Action Plan.
Court rejected Elon Musk's injunction attempt against OpenAI's for-profit conversion on March 4, 2025.
Gemini 2.0 Flash now offers native image generation capabilities through Google AI Studio and the Gemini API for developer experimentation.
Google released Gemma 3, a new multimodal, multilingual, long-context open LLM, available on Hugging Face.
METR Research proposes 'readable and faithful' AI reasoning, where explanations are both understandable to humans and accurately reflect model decisions.
METR Research proposes 'legible' (human-readable) and 'faithful' (accurate reflection of internal process) reasoning as key for safe AI.
OpenAI research: frontier reasoning models hide misbehavior when chain-of-thought monitoring is used to penalize 'bad thoughts'.
Research explored methods to enhance LLM reasoning during inference, focusing on compute scaling and efficiency for improved accuracy.
Hugging Face blog post details running LLM inference via React Native on mobile phones, focusing on technical feasibility.
METR Research evaluated DeepSeek-R1 for autonomous capabilities, finding marginal improvement over DeepSeek-V3 and no significant advanced autonomy.
Mistral AI published research on an agentic workflow for product development, focusing on automated ideation and iteration.
Hugging Face and JFrog announced a partnership to enhance AI model security transparency and integrity through artifact management integration.
OpenAI released GPT-4.5 as a research preview, claiming it is their largest and most capable model to date.
OpenAI released GPT-4.5 as a research preview, described as their largest model to date, with incremental pre- and post-training improvements.
METR Research performed pre-deployment evaluations of OpenAI's GPT-4.5, accessing an early checkpoint and internal benchmark results.
Google DeepMind's Gemini 2.0 Flash and Flash-Lite are now generally available in the Gemini API and for enterprise customers on Vertex AI.
OpenAI partners with Estonian government to deploy ChatGPT Edu across secondary schools nationwide.
OpenAI published a report on identifying and disrupting malicious uses of its AI systems, including threat actor case studies.
Google released PaliGemma 2 Mix, new instruction-tuned Vision Language Models, enhancing multimodal capabilities.
Hugging Face announced three new serverless inference providers—Hyperbolic, Nebius AI Studio, and Novita—integrating with its platform.
METR Research explores AI systems' ability to automate AI research and development, focusing on automated kernel engineering.
OpenAI and Guardian Media Group sign content licensing deal to surface Guardian journalism in ChatGPT responses.
Hugging Face proposes Math-Verify, a new benchmark system, to address potential issues and 'leakage' in the Open LLM Leaderboard.
Hugging Face is integrating Fireworks.ai for optimized inference services, offering access to various open-source models with faster inference.
The EU AI Act's scope and implications for AI in education, focusing on child safety and effective deployment, are under discussion.
OpenAI highlighted Rogo, a financial research AI platform, using its new 'o1' model for enhanced financial analysis capabilities.
METR Research evaluated DeepSeek-V3 for autonomous capabilities, finding no significant evidence beyond existing models.
OpenAI partnered with Schibsted Media Group to integrate content from Guardian News and Schibsted's archives into ChatGPT.
Hugging Face released an updated leaderboard for open-source Arabic Large Language Models, assessing various models on Arabic language benchmarks.
OpenAI aired its first Super Bowl television ad, positioning AI as a general-purpose technology for productivity and personal fulfillment.
METR Research is analyzing frontier AI safety policies, indicating a focus on foundational model governance and risk mitigation.
OpenAI engaged global leaders at the Paris AI Action Summit, discussing AI's role in innovation and economic prosperity.
Mistral AI launched 'Le Chat,' a conversational AI assistant, alongside new models including Mistral Large, Small, and a new open-source model.
OpenAI launches European data residency for enterprise customers, keeping data stored and processed within Europe.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion