To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval
Research analyzes scaling trade-offs between pretraining data scale and RAG retrieval store size across models up to 3B parameters.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research analyzes scaling trade-offs between pretraining data scale and RAG retrieval store size across models up to 3B parameters.
TraceSafe-Bench introduces a framework to benchmark LLM safety guardrails across multi-step, intermediate tool-calling execution traces.
Study replicates LLM package hallucination rates on frontier coding models, highlighting ongoing slopsquatting supply chain security risks.
Research reveals safety and trust benchmark scores drift significantly across minor updates within open-source LLM release lines.
Researchers introduced Poli-Bias, a counterfactual framework evaluating how LLMs alter legal and political reasoning based on country names.
Researchers identify systemic state-persistence failures in five major agentic workflow frameworks, proposing a machine-checkable resume contract.
Research identifies a critical security vulnerability in reasoning models where adversarial inputs starve safety monitors of token budget.
Researchers propose FitText, a memetic retrieval framework designed to help AI agents dynamically select tools from massive API ecosystems.
Researchers introduced LoopsBench, a long-horizon benchmark evaluating coding agents on sequential tasks structured as dependency DAGs.
Researchers proposed RoMeRL, a framework addressing state-space dispersion and misleading feedback loops in self-evolving LLM agent memories.
Researchers introduce GradCuit, a method enabling credit-assigned gradient flow to improve test-time latent reasoning in LLMs.
Academic paper demonstrates that LLMs struggle to suppress pretrained word meanings when system prompts redefine terms, causing silent failures.
Researchers use gradient-based training-data attribution to map which pretraining document regions support social versus STEM reasoning in OLMo3-7B.
Researchers introduce End-to-End Fairness Optimization (E2EFO), a framework integrating group fairness across prediction and decision stages.
New research explores 'any-order inference' for AI models, proposing masked diffusion models as a native approach for non-causal reasoning.
Academic research challenges the validity of AI benchmarks, proving that individual performance metrics cannot be reliably composed or extrapolated.
Researchers propose a pre-processing defense to close the 'decode gap' where encoded or imaged text bypasses Vision-Language Model safety guards.
Researchers introduce a fine-grained classification framework to detect and categorize inconsistency types in financial disclosure texts.
Researchers created GEMCo, a human-written proxy dataset of 86 German e-mail counseling conversations, validated against 124 real conversations.
New research identifies a critical flaw in LLM context attribution methods: they fail to distinguish between information retrieved from input context and knowledge encoded in model weights, producing unreliable scores.
Researchers propose unified static-dynamic pruning to improve LLM inference efficiency, addressing computational and memory bottlenecks in autoregressive decoding.
Research characterizes self-training (ST) in linear classifiers using pseudo-labels on Gaussian mixture data to understand generalization improvement.
Researchers propose Graph Wavelet Compressed Sensing (GWCS) for efficient, offline compression of graph signals to reduce data and training costs.
New research details cost accounting for exhaustive site-by-site interventions on neural network computational graphs, focusing on recomputation costs.
Research introduces a unified framework for Riemannian deep learning with reusable modules and manifold-specific architectures.
Research finds that vision-language models like CLIP primarily rely on class names for descriptions, not visual evidence, leading to poor zero-shot accuracy.
QuArch is presented as the first benchmark to evaluate LLM knowledge and reasoning in computer architecture, bridging software and hardware.
Researchers propose HindsightBench, a black-box audit protocol to detect parametric hindsight (leakage of future knowledge) in LLMs used for time-indexed decision tasks.
FlashRT is a research project exploring an agent-based harness to optimize deployment of real-time multimodal AI applications by managing model placement and parallelism.
Research introduces "Exact Network Surgery," a method to insert residual blocks into live neural networks while preserving function bit-exactly.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion