Algorithmic Compliance and Regulatory Loss in Digital Assets
ML-based AML systems in cryptocurrency show poor real-world performance due to temporal nonstationarity, despite strong static metrics.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
ML-based AML systems in cryptocurrency show poor real-world performance due to temporal nonstationarity, despite strong static metrics.
New research proposes a computationally efficient method for membership inference attacks (MIAs) on Diffusion Models (DMs) by analyzing predicted noise vectors.
Research introduces a unified taxonomy for categorizing Deep Learning-based Multivariate Time Series Anomaly Detection (MTSAD) methods.
Research explores methods for estimating rare, worst-case outputs from language models to improve safety evaluations beyond average behavior.
Research explores using LLMs as evaluators for information retrieval relevance, extending prior studies on LLM assessor effectiveness.
A research paper introduces the Consensus-Bottleneck Asset Pricing Model (CB-APM), a deep learning model for stock returns designed for interpretability-by-design through an analyst consensus bottleneck.
Research introduces a group matching score to address systematic underestimation of multimodal model capabilities in compositional reasoning benchmarks.
Research proposes Sovereign Agentic Loops (SAL) to decouple LLM reasoning from execution, mitigating safety risks in real-world systems.
Research identifies shared lexical task representations as a cause of LLM prompt sensitivity, comparing instruction-based and example-based prompting.
Research indicates that for 1-3B parameter models, execution feedback is more critical than complex pipeline topology for code generation.
Research explores post-training N:M activation pruning for LLMs, aiming for more efficient inference by dynamically compressing activations.
Research suggests learning rate decay in curriculum-based LLM pretraining wastes high-quality data, hindering performance gains.
Research paper models benchmark hacking in ML contests, showing how models are tuned to score highly without true generalization.
Researchers propose TS-Arena, a live forecasting platform for Time Series Foundation Models, to address train-test overlap risks in evaluation.
Research identifies classification model output label space as a privacy side-channel, demonstrating a concrete privacy attack despite Differential Privacy (DP) training.
Research proposes Utility-Aligned Embeddings (UAE) to enhance RAG dense retrieval by distilling LLM re-ranking utility, aiming for better precision and efficiency.
Researchers propose a formal definition for the "jailbreak oracle problem" to systematically assess LLM vulnerability to security bypasses.
A new framework, Sum-of-Checks, enhances auditability and reliability of Large Vision-Language Models for safety-critical tasks like surgical assessment.
Research proposes a statistical framework for evaluating multi-agent LLM systems, addressing reliability and error accumulation in safety-critical applications.
Research explores adversarial generation of Linux ELF malware using semantic-preserving transformations, addressing a gap in Windows PE-focused studies.
Research describes Stealth Pretraining Seeding (SPS), a new attack family embedding logic landmines in LLMs via poisoned web content during pretraining.
Research explores algorithms that highlight subsets of case-specific features for human decision-makers, rather than generating a single prediction.
Research paper proposes WassersteinGrad, a gradient-based method to explain autoregressive neural network predictions on dynamic physical fields.
PrivUn framework evaluates machine unlearning effectiveness in LLMs against privacy attacks, assessing direct retrieval and in-context recovery.
Research identifies 'background temperature' as a formal concept for hidden randomness in LLM outputs, even at T=0, due to implementation details.
Researchers propose "Kernel Contracts," a specification language for defining the expected behavior and correctness of ML kernels across diverse hardware.
Hugging Face blog post discusses using OpenAI's Privacy Filter for scalable web applications.
OpenAI highlights Choco's use of OpenAI APIs and AI agents to automate food distribution, increasing productivity and operational growth.
An article argues that optimizing for token count alone (tokenmaxxing) is not a sustainable AI strategy without considering fit and actual value.
Former AWS executive emphasizes organizational change and people management over technology for successful enterprise AI adoption.
OpenAI's Frontier Lab released a framework analyzing 921 occupations and 148 million US jobs for AI automation, reorganization, or growth potential.
DeepSeek released V4-Pro (1.6T total params, 49B active) and V4-Flash (284B total, 13B active), both 1M context Mixture-of-Experts with MIT license.
Latent Space claims OpenAI is developing GPT-5.5 and a 'Codex Superapp' to integrate agents for complex task execution.
Research estimates the value of additional recurrence in looped language models, proposing a new recurrence-equivalence exponent of 0.46.
Research proposes methods to measure and quantify environmental factors influencing LLM propensity for unsanctioned behavior, using Bayesian GLMs.
Research proposes new benchmarks for LLMs to assess genuine program execution understanding beyond surface-level code patterns or specific input prediction.
Research identifies regional cultural biases in LLMs, specifically an overrepresentation of Japanese culture in responses to cultural queries.
Research investigates LLMs and AI agents for automating the diagnosis and repair of computational research reproducibility failures due to code and environment issues.
Research claims prior work underestimates code generation bias by testing ML pipeline generation instead of simple if-statements.
Research introduces RedirectQA dataset to analyze LLM factual memorization beyond canonical entity names, focusing on how different surface forms affect recall.
Research introduces a 'prefix grammar transformation' to efficiently reduce prefix parsing to ordinary parsing, relevant for syntactically constrained LLM generation.
Research paper proposes method to detect and quantify opinion bias and 'sycophancy' in LLMs by observing responses to coercive prompts.
Research characterizes LLM behavior in whistleblower dilemmas, varying crime severity and relational closeness, evaluating moral judgment and predicted human actions.
Research claims LLM agent distillation leads to behavioral homogenization, making models share reasoning steps and failure modes from teacher models.
Research proposes inference-level mitigation for LLM fairness, addressing limitations of training-time methods in adaptiveness and computational cost.
Research identifies a new class of stealthy backdoor attacks against LLMs using natural language style triggers, avoiding explicit patterns.
RewardBench 2 introduces new benchmarks for evaluating reward models, which are critical for aligning LLMs with human preferences and safety.
Research benchmarks how LLM-based speech recognition systems' text priors affect demographic bias compared to traditional ASR architectures.
Research identifies prompt-induced hallucinations in large vision-language models, where prompts override visual input.
Research identifies novel 'function hijacking' attacks against agentic LLMs, exploiting vulnerabilities in external function calling mechanisms.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion