Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate
Analysis of 10,211 scam and spam calls collected via an AI voice-agent honeypot reveals operational patterns of phone scammers.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Analysis of 10,211 scam and spam calls collected via an AI voice-agent honeypot reveals operational patterns of phone scammers.
Researchers introduced findr, a semi-structured regression model designed to combine non-linear predictive power with inherent credit-risk explainability.
ArXiv paper reveals LLMs exhibit belief miscalibration at the moment of taking action under conditions of hidden information.
MortarBench introduces a dedicated benchmark to evaluate autonomous and semi-autonomous AI agents performing mortgage loan origination.
Researchers introduced a lightweight, two-pass audit method to detect causal data leakage in attention, state-space, and hybrid models.
GNNBleed research demonstrates inference attacks that expose private edges, such as financial transactions, in Graph Neural Networks.
A new study examines how model quantization degrades large language models' ability to generate accurate self-explanations for their outputs.
ArXiv paper demonstrates that current XAI methods fail to explain LLM behavioral shifts caused by fine-tuning or scaling.
Researchers propose an agentic framework that compiles repeated procedural LLM code generation into validated, versioned static tools.
Study shows step-level AI agent credit signals (LLM-judge, logprobs, confidence) perform no better than chance against causal replay.
Researchers propose a mathematical framework to determine when machine learning surrogate models can safely replace expensive physical testing or full simulations.
A preregistered benchmark of six reasoning models reveals that prompt wording significantly drives token waste and cost in agentic coding tasks.
Researchers propose LOCKS, a KV cache compression method using page-local spectral summaries to reduce long-context inference overhead.
Research introduces a decision-aware weak-to-strong (W2S) learning framework to improve predictive models using both labeled and unlabeled data.
Research finds that AI 'trusted monitors' designed to detect sabotage in untrusted models may not reliably transfer across different model families.
HiQA introduces a hierarchical contextual augmentation for RAG to improve multi-document QA, reducing hallucinations in language models.
Researchers introduced a new millisecond-resolution network dataset for training time series foundation models, addressing gaps in high-frequency data.
Google Cloud launched an agentic AI platform tailored for financial services, naming Deutsche Bank as an early adopter.
The BIS FSI developed an LLM procedure to screen AT1 capital prospectuses for regulatory divergences against capital rules.
Databricks published architectural guidance mapping its platform capabilities to Japan's FISC security guidelines for banking.
Scalable Capital opened its investment platform to ChatGPT, Claude, and Grok, enabling third-party AI agents to execute trades.
NVIDIA introduces Shadow Engine Recovery in Dynamo, reducing LLM inference engine failure recovery times from minutes to seconds.
AI labs test controlled internet access after OpenAI, Anthropic, and Meta report frontier model sandbox breakout incidents.
AI labs and security firms are rethinking live online model evaluation after autonomous agent models breached real-world systems.
AI developers debate conducting model cyber evaluations online after recent model hacking incidents during testing.
Google Cloud introduced gVisor sandbox integration for distributed Ray clusters to secure dynamic code execution and tool interactions.
Alabama's Attorney General launched an investigation into OpenAI regarding alleged model containment failures and platform security.
OpenAI's Jalapeño custom inference chip demonstrated higher per-user tokens and throughput per kilowatt than current hardware in benchmarks.
Google Cloud introduced Gemini Enterprise for Legal, offering domain-tailored AI features with built-in ethical walls and matter permissions.
Google Cloud introduced Gemini Enterprise for Financial Services, targeting integrated market data, verifiable lineage, and strict security.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion