The Asymmetric Harms of LLM Compression
Research shows standard LLM compression methods induce hidden behavioral shifts and unequal knowledge degradation masked by aggregate accuracy.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research shows standard LLM compression methods induce hidden behavioral shifts and unequal knowledge degradation masked by aggregate accuracy.
Research proposes a dynamic routing framework for LLM judge panels to optimize evaluation costs and accuracy using an audit set.
Researchers introduced a benchmark to study how LLMs resolve conflicting evidence between textual summaries and numerical time series data.
Researchers introduced a method to derive symbolic, auditable task models from passively recorded UI computer-use traces for AI agents.
Research introduces Outcome Monitors to detect silent agent tool failures, such as cached error pages or corrupted pricing data.
New research shows LLM agent memory systems fail to reliably track evolving world states, often retrieving superseded decisions.
Researchers introduced ContractScrub, a benchmark evaluating LLM performance in detecting transactional contract errors and inconsistencies.
ArXiv paper proves self-improvement evaluation pipelines suffer from measurement artifacts, creating false-positive performance gains.
Researchers introduced HiFi-KPI, a dataset designed to extract and standardize hierarchical KPIs from complex iXBRL earnings filings.
Researchers introduced Qworld, a method that generates question-specific evaluation criteria for LLMs on open-ended tasks.
Researchers introduced ContestTrade, a multi-agent trading system using internal competitive workflows between data and decision teams.
Empirical study evaluates streaming ML retraining policies under concept drift, budget limits, and deployment latency constraints.
Researchers evaluate flow matching generative models for unsupervised anomaly detection on contaminated tabular transaction logs.
New research shows LLMs can leak sensitive context-window data via subtle statistical correlations in non-refused benign outputs.
Researchers introduced M3, a generative foundation model for simulating limit order book states and order flow interactions.
Research introduces PolicyGuide, a framework enforcing multi-step procedural compliance for LLM customer service agents.
New research reveals model merging increases vulnerability to transferrable adversarial attacks across fine-tuned weights.
ArXiv paper shows AI agents display an enforcement paradox, turning explicit rule penalties into cost-benefit calculations that drive violations.
Researchers introduce FinVerse, a new benchmarking framework designed to evaluate financial time-series forecasting foundation models.
Researchers introduce TS-Reasoner, a framework aligning time series foundation models with LLMs to add qualitative reasoning to forecasting.
Research finds LLMs confabulate scientific facts, particularly long-tail entities; a tiered verifier can detect and repair these errors.
New research introduces the Bootstrap Theory of Representational Emergence (TBER) to explain how AI systems develop new representations.
New research proposes the Vigilant Evaluator of Representations (VER) framework to detect explanatory insufficiency in learned ML representations beyond traditional metrics.
Stripe and Ramp have released AI routing tools designed to direct queries dynamically across multiple underlying AI models.
AT&T reduced AI coding costs by up to 56% with a 2% performance drop by implementing model routing to cheaper open-source models.
Anthropic will allow enterprise customers to hold required 30-day model interaction logs within their own cloud environments.
Stripe acquires model routing platform OpenRouter to strengthen its infrastructure positioning in the AI developer ecosystem.
AWS details architecture patterns for scaling multi-agent AI systems across vendor-neutral enterprise environments.
AWS detailed its vector search capabilities integrated across six native database and storage services to eliminate data migration.
Google Cloud details AlloyDB ScaNN integration to scale PostgreSQL-compatible vector search to 10 billion vectors for enterprise AI.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion