Brit mathematician lets AI agent loose with credit card – cue password leaks, CAPTCHA chaos and more
A mathematician's AI agent, given a credit card, leaked passwords and struggled with CAPTCHAs, demonstrating agentic tech risks.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A mathematician's AI agent, given a credit card, leaked passwords and struggled with CAPTCHAs, demonstrating agentic tech risks.
OpenAI released MRC (Multipath Reliable Connection), a new supercomputer networking protocol, via OCP to enhance resilience and performance for large AI training.
OpenAI announced GPT-5.5 Instant, updating ChatGPT's default model with claimed improvements in accuracy, reduced hallucinations, and personalization.
OpenAI published a 'System Card' for an unreleased model, GPT-5.5 Instant, detailing internal safety evaluations and intended capabilities.
OpenAI launched a European Youth Safety Blueprint and EMEA Youth & Wellbeing Grants to promote safe and responsible AI use for young people.
OpenAI introduced a self-serve Ads Manager for ChatGPT, featuring CPC bidding and enhanced measurement tools, with privacy safeguards.
Microsoft fixed VS Code's Git extension after it automatically credited Copilot as a co-author for human-written code, raising developer concerns.
OpenAI and PwC announced a partnership to help enterprises automate finance workflows, improve forecasting, and modernize CFO functions using AI agents.
Agentic AI models will fundamentally change cloud storage architecture demands due to their unique inference patterns, unlike traditional workloads.
Expert commentary suggests AI systems are approaching recursive self-improvement capabilities.
TeamViewer ONE is promoting agentic AI for IT support to proactively resolve issues, moving from reactive troubleshooting to automated prevention.
OpenAI details its optimized WebRTC stack for real-time, low-latency Voice AI with global scale and conversational turn-taking.
AI chip startups are finding new opportunities in the inference market, challenging Nvidia's dominance in a disaggregated AI landscape.
Eugene Yan outlines five principles for building and scaling AI systems, focusing on context as infrastructure and verification for autonomy.
Report discusses local AI coding agents as an alternative to usage-based LLM pricing, avoiding token limits and promoting 'vibe coding'.
Forrester argues that CIOs must manage the systematic risk of AI-generated software amid concerns over “systematic failure at scale.”
Mozilla criticizes Google for integrating a Prompt API directly into Chrome, expressing concerns that it will reduce the openness of the web.
SAP user groups express concern over new API policy, fearing it will hinder adoption of AI and other innovations integrated with SAP systems.
Anthropic reportedly generates more revenue than OpenAI from its LLM offerings by focusing on enterprise customers with higher-value use cases.
Amazon's custom AI chip business, including Trainium, has grown to an estimated $20 billion, signifying increased internal and external adoption.
OpenAI detailed the root cause and mitigation for 'goblin' outputs in GPT-5, attributing personality-driven quirks to specific training data.
Researchers demonstrated that LLMs can be easily poisoned through a $12 domain registration and one Wikipedia edit to alter factual recall.
Hugging Face details the training methodology and architectural choices behind IBM's Granite 4.1 series of LLMs, focusing on pre-training data.
OpenAI announces 'Stargate' initiative, a massive compute infrastructure project to support AGI development and meet future AI demand.
AWS engineers internally advise against AI shortcuts and advocate for human review, contrasting with public hype.
OpenAI published a five-part action plan for cybersecurity in the 'Intelligence Age,' emphasizing AI-powered defense and critical system protection.
DeepInfra is now available as an inference provider on Hugging Face, enabling easier deployment and scaling of open-source models.
The FCA announced the second cohort for its AI Live Testing initiative, including Barclays, Lloyds (Scottish Widows), and UBS.
FCA's Jessica Rusu highlighted agentic commerce and Open Finance as key innovation drivers, announcing an expansion of their AI Lab.
At AI Dev 26 x SF, software developers discussed the impact of AI on their roles and the future of software development.
OpenAI models are now available in limited preview on AWS Bedrock, expanding distribution beyond Microsoft Azure and direct API access.
A new chatbot, Talkie, is trained only on data up to 1930 to explore how AI processes historical information and 'thinks.'
IBM's AI coding assistant, "Bob," is now generally available after internal testing with 80,000 users, including mainframe development.
Amazon introduced a new 'Copilot' AI assistant designed to integrate across various enterprise applications, emphasizing always-on context.
NVIDIA introduces Nemotron 3 Nano Omni, a multimodal LLM for long-context document, audio, and video agents.
Tenstorrent launched new RISC-V based AI servers, Galaxy Blackhole, featuring 32 accelerators in a 6U chassis for $110,000.
Brussels' Digital Markets Act enforcers demand Google grant rival AI assistants the same deep Android device access as Gemini.
Report highlights that G-SIB executives underestimated the difficulty of swapping AI models, leading to significant vendor lock-in.
AI adoption is driving revenue deflation in India's tech services sector, with project values decreasing despite stable headcounts.
MEMCoder research introduces a multi-dimensional evolving memory system for LLMs to improve code generation using private enterprise libraries.
Research identifies a flaw in audio-language model evaluation: models can achieve high scores on audio benchmarks using text priors, not true audio understanding.
AdaComp is a new context compression method for RAG that uses an adaptive predictor to extract relevant sentences, aiming to reduce noise and cost.
Researchers introduced an N-gram Coverage Attack, a membership inference method effective against API-only LLMs like GPT-4, without hidden state access.
Research identifies prompt underspecification as a key source of LLM instability, leading to significant performance degradation when prompts or models change.
Research proposes an evaluation framework for highlight explanations, aimed at showing which context pieces LMs use to generate responses.
Research paper SWE-QA introduces a new benchmark for evaluating LLMs' ability to answer complex, repository-level code questions beyond simple snippets.
Research demonstrates a new 'intention deception' method for jailbreaking frontier LLMs, exploiting brittleness in current safety alignment.
Research introduces SpeechLLMs for direct speech processing, questioning if it improves speech-to-text translation quality over cascaded methods.
Research evaluated agentic LLMs on synthesizing longitudinal multiple myeloma patient records against expert clinical consensus for treatment decisions.
Research systematically benchmarks context utilization techniques (CMTs) for language models, addressing issues of ignored or irrelevant information.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion