Cursor-Opus agent snuffs out startup’s production database
An AI agent using Cursor and Opus deleted a startup's production database during an automated task, though data was recovered.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
An AI agent using Cursor and Opus deleted a startup's production database during an automated task, though data was recovered.
OpenAI's long-standing AGI clause with Microsoft, which would have nullified commercial IP rights upon AGI achievement, has been removed.
South Africa retracted an AI policy document after an AI-assisted drafting process resulted in fabricated citations and non-existent legal references.
Executives from Citi, Home Depot, and Capcom shared lessons learned from early AI agent deployments in financial services, retail, and gaming.
OpenAI is reportedly free to partner with other cloud providers beyond Microsoft's Azure until 2032, ending prior exclusivity.
OpenAI's ChatGPT Enterprise and API achieve FedRAMP Moderate authorization, clearing secure AI adoption for U.S. federal agencies.
OpenAI and Microsoft announced an amended agreement clarifying their partnership terms to support continued AI innovation and scale.
BIS paper analyzes Samsung's decision to not patent dual SIM technology in India, examining implications for innovation diffusion.
WorkHQ claims to offer scalable agentic automation benefits, moving beyond experimentation in the enterprise AI space.
Google DeepMind partners with the Republic of Korea to advance scientific research using frontier AI models.
New research proposes Logit-Balanced Vocabulary Partitioning (SSG) to improve LLM watermarking, specifically KGW, in low-entropy text like code.
Research explores methods for LLM-generated business idea evaluation, focusing on whether automatic judges should aggregate expert consensus or model individual evaluators given disagreement.
Research proposes using embedding models to improve probabilistic race prediction, addressing limitations of traditional Census-based methods like BISG for uncommon surnames.
Research proposes RouteLMT, a learned routing method for hybrid LLM translation systems, balancing cost and quality over heuristic approaches.
Research finds AI writing assistance distorts perceived writer persona, affecting beliefs, personality, and identity across 29 social dimensions.
Research indicates standard RL from Verifiable Rewards (RLVR) may not guarantee a model's stated chain-of-thought reasoning is causally important to its answer.
Researchers developed a highly efficient RAG system for Ukrainian document Q&A, achieving 2nd place in the UNLP 2026 Shared Task.
Research finds leading LLMs (Claude Sonnet 4.5, GPT-5.4, Gemini 2.5 Flash) exhibit individualism-collectivism bias in advice, varying by country and language.
Research explores how LLMs resolve conflicts between internal knowledge, user assertions, and retrieved document content in RAG and chat systems.
LLMs often determine final answers early, with subsequent chain-of-thought tokens serving as post-decision explanations, increasing inference cost.
Research finds LLM rewriting significantly alters personal narratives, reducing distinct linguistic markers across 13 stylistic measures.
LLM-generated narratives perpetuate representational harms against global majority nationalities, highlighting bias risks in enterprise applications.
Research systematically analyzes token consumption in AI agents during coding tasks, identifying cost drivers and exploring prediction methods.
Research finds LLMs are highly persuasive in everyday conversations, outperforming humans, and users consult them for major life decisions.
Research finds LLMs' advice defaults often conflict with community-endorsed moral orders, highlighting alignment challenges in prescriptive tasks.
Research proposes a new method, "Behavioral Canaries," to audit if private retrieved contexts are illicitly used in LLM RL fine-tuning.
Research proposes a structured reasoning framework for scalable question answering over long document sets, addressing LLM context window limits.
Research evaluates methods for selecting optimal query variants in RAG pipelines prior to full retrieval, aiming to reduce computational cost.
Research finds multilingual LLMs can improve question answering by changing input query language, introducing the concept of Language Specific Knowledge (LSK).
Research proposes automated methods for evaluating the robustness of LLMs in mathematical reasoning, addressing limitations of current manual evaluations.
Research investigates methods for generating closed-ended survey responses using LLMs to simulate human survey participants in-silico, aiming for a standard practice.
Research explores singular value decomposition compression and tiling for efficient LLM inference on AWS Trainium accelerators.
NiuTrans.LMT research identifies a performance degradation mode in multilingual machine translation LLMs when fine-tuned symmetrically on pivot data.
Research identifies system-mediated attention imbalances, not just image attention, as a key factor in vision-language model hallucinations.
Research introduces 'source-modality monitoring' in multimodal models, evaluating their ability to track input origin for information binding.
Research finds LLMs struggle to detect culture-specific health misinformation, using cow urine discourse in India as a case study.
Research indicates Diffusion-based LLMs (dLLMs) like LLaDA and Dream underperform auto-regressive models for agentic workflows, despite claims of latency reduction.
Research systematically studies optimal LoRA adapter placement in hybrid language models (attention + recurrent components) for fine-tuning efficiency.
Research explores online distributional regression for large-scale streaming data, focusing on learning conditional heteroskedasticity in probabilistic forecasting.
Researchers propose Controllable Alignment Prompting (CAP) for LLM unlearning, addressing cost and access issues for closed-source models.
Research presents a multi-layered methodology to accelerate multimodal foundation models through hardware and software co-design and optimization.
MCAP is a new research method to profile LLM layers at deployment time, optimizing memory use for inference across heterogeneous hardware.
Research critiques common Shapley-based XAI evaluation methods, showing fragmented approaches lack human utility verification in high-stakes contexts.
Research explores feature attribution methods for Supervised Contrastive Learning (SCL) models, an alternative to cross-entropy for classification.
Researchers extended Neural Activation Coverage (NAC) for uncertainty estimation in regression models, claiming superior results over Monte-Carlo Dropout.
Research claims foundation models outperform dataset-specific ML for energy time series forecasting, suggesting broad applicability.
Research introduces Atlas-Alignment, a method to make interpretability techniques transferable across language models, reducing the cost of model-specific interpretation.
Research investigates how LLMs detect and correct their own errors using internal confidence signals, distinct from first-order self-evaluation.
Research applies persistent homology to characterize how adversarial inputs reshape LLM internal representation spaces, moving beyond linear interpretability.
Research identifies universal adversarial perturbations that compromise modern behavior cloning policies, a common method for training AI from demonstrations.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion