Introducing Gemini 3.7 Flash
Google DeepMind announced the release of Gemini 3.7 Flash, expanding its lightweight, high-speed model lineup.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Google DeepMind announced the release of Gemini 3.7 Flash, expanding its lightweight, high-speed model lineup.
Google Cloud introduced BigQuery Graph with measures in preview to ground enterprise agentic workloads in structured graph data.
AWS released a walkthrough using OpenTelemetry to route agent traces, metrics, and token usage from on-prem, GCP, and Azure to Bedrock.
AWS detailed a reference architecture using Amazon Bedrock AgentCore Browser Tool to automate legacy web apps via isolated sessions.
Anthropic is reportedly in talks to acquire chip-optimization startup Decart AI for $6 billion to reduce compute costs.
AWS released a reference architecture and sample code for multi-agent M&A due diligence using Amazon Bedrock AgentCore.
AWS released Amazon Quick extensions for Microsoft 365, integrating agentic document editing directly into Word, Excel, and Outlook.
The White House plans to extend its voluntary pre-release cybersecurity testing framework to cover powerful open-weight AI models.
Microsoft is consolidating its consumer and enterprise Copilot applications into a single unified interface supporting work and personal accounts.
Mindgard secured $30 million in Series A funding to expand its automated AI security testing and vulnerability assessment platform.
DeepSeek is steeply increasing API prices for its V4 models, closing the cost gap with established major AI rivals.
AI debt accounts for up to 30% of net new credit issuance, exposing bank balance sheets to GPU residual value risks.
OpenAI previews Ultrafast API tier for GPT-5.6 Sol, leveraging Cerebras hardware to deliver up to 750 output tokens per second.
Bank of England research analyzes financial stability risks and market spillovers if US big-tech AI earnings fail to meet valuations.
Standard Chartered argues long-dated capital overcommitment to AI infrastructure poses a greater structural risk to banks than low returns.
Researchers released Backtrader-Bench, an evaluation framework using dynamic MCQs to benchmark LLM coding agents in algorithmic trading.
New benchmark COMPINT reveals LLM context compaction silently drops persistent session constraints and safety instructions in multi-turn sessions.
Research shows LLMs retain latent knowledge after unlearning, which can recover during subsequent fine-tuning or continued training.
Research shows adapting LLMs to specific group preferences increases sycophancy, causing models to prioritize agreement over factual truth.
Research shows semantic rephrasing of benchmark prompts routinely flips LLM output correctness, exposing high evaluation instability.
LabelFusion-TS fuses text models and market time series to classify Federal Reserve communications into hawkish, dovish, or neutral stance.
New research benchmarks the serving costs of agentic memory frameworks against standard context window strategies over long conversations.
New arXiv research benchmarks trustworthiness in small language models, comparing natively trained SLMs against compressed variants.
Researchers released FrontierFinance, a benchmark evaluating AI agents on complex, open-ended investment research workflows.
ArXiv research analyzes performance trade-offs in gist-based context compression, showing where summarization impairs agent memory retrieval.
Research demonstrates China-origin vision-language models systematically reframe sensitive visual inputs rather than issuing explicit refusals.
ToolHazard creates scalable adversarial environments to evaluate LLM agent vulnerability to indirect prompt injections.
Research shows LLM performance rankings flip when generation token budgets vary, with 3–19% of tasks showing accuracy drops at higher caps.
Paper proposes integrating LLM-derived aleatoric and epistemic uncertainty directly into portfolio covariance matrices for small-cap trading.
A systematic study shows LLMs struggle to accurately translate natural language constraints into formal solver code for complex problems.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion