Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
A systematic arXiv scaling study evaluates the accuracy-cost curve of lexical, dense, graph, and agentic RAG across corpus sizes up to 512k docs.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A systematic arXiv scaling study evaluates the accuracy-cost curve of lexical, dense, graph, and agentic RAG across corpus sizes up to 512k docs.
Researchers introduced WikiLoop, a framework that jointly optimizes knowledge-base construction and agent navigation using downstream feedback.
Researchers analyze filesystem-based memory for LLM agents, evaluating how agents manage growing directory structures of markdown files.
Researchers propose shifting from extracting concepts post-training (dictionary learning) to designing explicit conceptual structures in LLMs.
Research shows LLMs lack robustness when processing non-canonical tokenizations, degrading performance inconsistently across different languages.
Researchers introduce SERPO, a method enabling language models to self-evolve at inference time for open-ended generation tasks.
Researchers introduced paired prompts to evaluate how well language models match diagnostic evidence to specific causal questions.
Researchers introduced CreditCardQA, a dataset of 1,800 questions based on real credit card agreements to test LLM numerical reasoning.
Academic paper "OptimismBench" identifies systematic optimistic bias and non-additive probability estimation errors in LLM decision judgments.
Researchers introduced the Stereotypes-to-Decisions framework to evaluate regional bias in LLMs from abstract views to concrete choices.
Pangram Labs released a technical report for Pangram 4, claiming a 0.9916 AUROC and a 0.0041% false positive rate for AI text detection.
Mercor and Ramp introduced APEX-Accounting, a private benchmark of 160 multi-file tasks designed to evaluate LLMs on accounting workflows.
Researchers analyzed GPT-4-Turbo to identify and measure implicit bias toward people with intellectual disabilities in LLM-generated stories.
Researchers introduce GPT-Red, an automated self-play red-teaming agent designed to find prompt injection vulnerabilities in frontier models.
Researchers find RL-trained reasoning models develop superior internal representational quality over SFT models for math problem-solving.
Research identifies 'trust inflation' in LLM evaluation, where aggregating weak and strong metrics masks vulnerabilities in model performance.
Researchers identify a 'confounder trap' in causal inference from text, where representations learned to control for confounding encode treatment.
Researchers proposed IRIS, a framework that extracts dynamic user personas from implicit interaction streams rather than explicit feedback.
Researchers demonstrate that audio-capable LLMs can be jailbroken solely through variations in vocal delivery (prosody) with fixed transcripts.
Researchers introduce SecRespond, a benchmark specifically designed to evaluate LLM agents in real-world post-compromise incident response tasks.
Research shows frontier multimodal models generate structured, biased confabulations based on demographic descriptors when input images are missing.
Researchers introduced Setoka, a new benchmark designed to evaluate personalized AI agents on retrieving and inferring user characteristics.
Researchers propose an on-policy distillation routing method to protect fine-tuned LLMs from malicious safety-realignment bypasses.
Researchers define 'linguistic monoculture,' showing that widespread LLM-assisted drafting reduces population-level variations in text.
Researchers introduce MindForge, a framework training small language models in whole-life-cycle software generation via synthetic data.
Researchers introduce SpecFirst, an agentic framework designed to improve program synthesis from scratch using behavioral exploration.
A study of Gemini 2.0 Flash and ChatGPT-4o finds significant diagnostic inconsistency under rephrased prompts and irrelevant content.
Researchers introduce ML2B, a benchmark of 35 Kaggle competitions in 14 languages to evaluate LLMs on cross-lingual ML pipeline generation.
Researchers propose ARC-Encoder, a context compression technique that works without modifying or fine-tuning the target LLM architecture.
Researchers analyzed how context alters truth representation vectors inside LLM activations, finding context deforms internal truth geometry.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion