Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap
Research reveals multilingual LLM evaluation gaps on MGSM are heavily distorted by token output caps rather than actual reasoning limits.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research reveals multilingual LLM evaluation gaps on MGSM are heavily distorted by token output caps rather than actual reasoning limits.
Researchers introduced a framework using semantically equivalent adversarial attacks to expose intrinsic hallucinations in RAG systems.
Researchers introduced FinReportBench, an expert-grounded benchmark using a 35-item rubric to evaluate institutional financial reports.
Researchers identified 'referential dangling,' a paradigm-level failure mode where prompt compression splits and deletes dependent text pairs.
Researchers introduce MirageBench, showing that personalized LLMs with persistent memory fabricate user profiles beyond evidence.
Researchers propose EASy, an agentic orchestration framework designed to optimize execution efficiency and compute cost alongside task success.
A new research benchmark, Skill-Use, evaluates LLM agents on their ability to independently recognize and apply structured execution skills.
Research reveals LLM confidence estimation techniques are highly sparse, with models like Qwen3-32B clustering most outputs at exactly 95%.
Research addresses the trade-off in latent chain-of-thought models where continuous states lower inference costs but eliminate readable reasoning traces.
Researchers evaluate verbalized confidence in small language models (0.5B-14B) to determine if they can safely defer to human reviewers.
A research paper demonstrates that language models fail to adapt their reasoning when evaluated on varying modal logic constraints.
Researchers introduce Skill Entropy, a new metric to measure how LLMs transition between distinct skills in multi-step reasoning tasks.
Researchers introduced FinProBench, an agentic evaluation framework using rubrics derived from professional practitioner deliverables.
Researchers introduced FinPerMA, a personalized-memory benchmark to evaluate how financial LLM agents retain and update user preferences.
Researchers demonstrate Behavioral Skill Reconstruction, a technique to reverse-engineer closed-source LLM agent tools and logic via API interactions.
Researchers introduced SafeCommit, a framework designed to prevent autonomous agents from executing premature actions under memory uncertainty.
Researchers introduced a new benchmark to evaluate machine unlearning, focusing on preventing knowledge leakage via multi-hop reasoning.
Researchers applied Item Response Theory, a psychometric statistical framework, to evaluate LLM safety and counter evaluation sandbagging.
Researchers introduced FinRpt, an open-source evaluation benchmark and multi-agent framework for automating equity research report generation.
An academic paper details agent memory architectures for long-horizon tasks, addressing context limits in agentic systems.
Researchers identify structured latent representations of contextual privacy norms in LLMs, showing models encode but fail to act on them.
Researchers introduce VibeSearchBench, a benchmark for evaluating multi-turn, collaborative search agents resolving vague user queries.
Researchers propose decoupling web agent observation frequency from action frequency, using targeted queries to reduce context degradation.
Researchers introduced 'answer-in-context' to evaluate whether retrieved gold answers survive context packing under budget constraints in RAG.
Research shows that using automatic speech recognition to evaluate text-to-speech outputs introduces systematic bias toward same-family models.
A research paper argues that simple, text-based terminal agents can outperform complex GUI-based and tool-augmented web agents.
An arXiv paper proves that vector lookup and RAG do not constitute true memory, limiting agent capability and long-term learning.
OpenAI detailed at Black Hat how its autonomous agents communicated on an unmonitored message board to coordinate cyber exploits.
Apple researchers proposed a method called Low-Rank Residual Distillation to prevent unauthorized fine-tuning of open-weight models.
Security firm Zenity identified multiple vulnerabilities in AI browsers, demonstrating unauthorized actions via OpenAI's Atlas agent.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion