Who Checks the Citations? Benchmarking Legal Hallucination Detection
A research study analyzes over 1,000 court filings containing fabricated AI citations, demonstrating that LLM hallucination rates in legal contexts remain a persistent risk.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A research study analyzes over 1,000 court filings containing fabricated AI citations, demonstrating that LLM hallucination rates in legal contexts remain a persistent risk.
Researchers evaluate whether internal-state LLM hallucination detection signals generalize across different languages and domains.
A study of 14 reasoning models reveals that test-time compute scaling fails to reduce factual hallucinations in knowledge-intensive tasks.
Jane Street leads a $2 billion investment in Australian data center operator Firmus Technologies, securing regional AI compute capacity.
Goldman Sachs Research projects global AI investment will exceed $1 trillion by 2026, driven by adjusted hyperscaler capex forecasts.
NatWest Group deployed an AI platform named Serene to detect early signs of financial vulnerability and customer distress.
Apple ML Research published a paper comparing the performance, latency, and arithmetic intensity of diffusion versus autoregressive language models.
Apple ML Research introduced Arbitrage, a speculative decoding technique that reduces reasoning model inference latency and computational cost.
Apple ML Research demonstrates scaling categorical flow matching for discrete data, offering an alternative to autoregressive language models.
LangChain released implementation guidance for establishing user-identity-linked authentication and authorization boundaries for active AI agents.
LangChain published a framework for evaluating agentic AI, covering dataset construction, grader design, and production readiness.
LangChain's blog outlines the distinction between agentic software frameworks, runtimes, and evaluation harnesses for development.
LangChain 1.0 launches Middleware to grant developers control over context engineering, model calls, and tool execution for agents.
LangChain outlines the necessity of explicit user feedback loops within agent observability frameworks to enable continuous learning.
LangChain released a technical guide on using observability tools to trace, debug, and evaluate multi-step AI agent reasoning paths.
Google, Amazon, and Microsoft back Agent Plugins 1.0.0, a standard directory spec for packaging Agent Skills and MCP servers.
The Model Context Protocol specification has been updated to a fully stateless core, enabling cloud-native scaling and serverless routing.
Google Cloud API Gateway launches Public Preview of a serverless model routing feature supporting Gemini, Claude, and OpenAI endpoints.
Google outlines infrastructure patterns for real-time AI agents using session-aware load balancing to manage stateful, bidirectional streams.
Google announced the general availability of its agent evaluation service, featuring over 20 metrics and LLM-as-a-judge capabilities.
Google released an open-source TPU microbenchmark suite to diagnose performance bottlenecks across network, compute, and memory.
Google released Tunix, a JAX-native library designed to optimize TPU throughput when training multi-turn, tool-using LLM reasoning agents.
Google proposed a modular prompt transpilation framework that treats system instructions as validated build artifacts in CI/CD pipelines.
Google Cloud integrates Parallel Web Systems search into Gemini, enabling developers to ground agent workflows in real-time web data.
Google engineers optimized Qwen 3.5-397B on Ironwood TPUs using JAX and a hybrid parallel topology, gaining 4.7x prefill speedups.
Google's JAX ecosystem introduces elastic training via Pathways, preventing multi-node training crashes by replacing only failed workers.
Google released the Genkit Agents API, an open-source framework for building multi-agent systems with managed state persistence.
Alibaba released its latest frontier AI model, aiming to compete with US tech giants in cloud and model capabilities.
AWS released a reference architecture using Amazon Bedrock AgentCore and Strands Agents SDK to automate financial complaint classification.
Google consolidates its AI units under Silicon Valley leadership, shifting focus from deep scientific research to commercial product delivery.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion