When Rule Violations Are Rare: Chimera Training for Logical Anomaly Detection
Research proposes "Chimera Training" for anomaly detection using logical rules on learned concepts, addressing rare rule violations.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research proposes "Chimera Training" for anomaly detection using logical rules on learned concepts, addressing rare rule violations.
Research explores using sparse autoencoders to interpret internal representations of a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR).
Researchers propose a method for LLM agents to retrieve sets of external tools simultaneously, rather than in isolation, by predicting hyperedges.
Research introduces a differentially private permutation test framework addressing privacy concerns in sensitive data analysis without significant statistical efficiency loss.
Research paper explores learning controlled stochastic differential equations (SDEs) from trajectory data to estimate drift and diffusion coefficients.
Research proposes an extended framework for financial volatility forecasting by incorporating a broader range of realized volatility measures.
New research introduces a method for clean-image backdoor attacks that maintains high clean accuracy while increasing attack potency, improving stealth.
New research proposes a physics-informed neural network for more accurate cross-covariance forecasting in financial markets, addressing limitations of traditional shrinkage methods.
Research indicates AI development in climate information exacerbates global inequality due to biased data and unrepresentative validation.
Researchers propose activation watermarking for LLM monitoring to detect misuse and prevent adaptive attackers from evading detection in deployed models.
REAP introduces an automatic method for curating coding agent benchmarks from interactive production usage, aiming to improve evaluation speed and fidelity.
Research reveals LLMs generate correlated fictional names, like 'Elena Vasquez' and 'Marcus Chen,' across AI-generated documents.
TLA-Prover, a 20B-parameter model, significantly improves verifiable TLA+ specification synthesis, achieving higher semantic model-check rates than other LLMs.
Researchers propose a validation methodology using Synthetic Customer Agents as digital twins to automate LLM chatbot testing in banking.
Researchers introduce V-Steer, a training-free inference-time method to enforce prompt hierarchy by editing value vectors in LLM cache.
Researchers introduced AgentGUI, a locally hosted interface designed to observe, log, and steer concurrent, long-running AI agent sessions.
Researchers developed an evaluation framework identifying systematic failure modes when using LLMs to simulate human survey respondents.
An academic study finds coding agents boost developer productivity but significantly degrade their understanding of the code they generate.
Research shows fine-tuning LLMs on narrow flaws causes broad behavioral misalignment that mimics systemic shifts in personality traits.
Researchers introduced ForgetBench, a benchmark evaluating how large language models retain or forget parametric knowledge during repeated updates.
A systematic arXiv scaling study evaluates the accuracy-cost curve of lexical, dense, graph, and agentic RAG across corpus sizes up to 512k docs.
Researchers introduced WikiLoop, a framework that jointly optimizes knowledge-base construction and agent navigation using downstream feedback.
Researchers analyze filesystem-based memory for LLM agents, evaluating how agents manage growing directory structures of markdown files.
Researchers propose shifting from extracting concepts post-training (dictionary learning) to designing explicit conceptual structures in LLMs.
Research shows LLMs lack robustness when processing non-canonical tokenizations, degrading performance inconsistently across different languages.
Researchers introduce SERPO, a method enabling language models to self-evolve at inference time for open-ended generation tasks.
Researchers introduced paired prompts to evaluate how well language models match diagnostic evidence to specific causal questions.
Researchers introduced CreditCardQA, a dataset of 1,800 questions based on real credit card agreements to test LLM numerical reasoning.
Academic paper "OptimismBench" identifies systematic optimistic bias and non-additive probability estimation errors in LLM decision judgments.
Researchers introduced the Stereotypes-to-Decisions framework to evaluate regional bias in LLMs from abstract views to concrete choices.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion