Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
New research proposes a method for evidence attribution in visual document understanding without needing explicit coordinate or region labels.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
New research proposes a method for evidence attribution in visual document understanding without needing explicit coordinate or region labels.
Research isolates performance factors in language model-based entity matching, distinguishing architecture from model backbone and size effects.
Research finds visually grounding a bilingual LSTM with images and captions in English and Spanish improves semantic understanding within and across languages.
Research indicates no optimal language set exists for multilingual instruction tuning, challenging the assumption that linguistically diverse sets universally yield better models.
Researchers introduced LibMoE, a unified framework for benchmarking Mixture of Experts (MoE) architectures in large language models due to high training/evaluation costs.
Researchers propose PReSS, an automated black-box framework to evaluate LLM political stance stability across topics, extending traditional bias classification.
New research proposes "Flick," a few-label text classification method using K-Aware Intermediate Learning to reduce reliance on extensive labeled data, especially for low-resource languages.
Researchers introduced LEDOM, a purely right-to-left autoregressive language model (2B/7B parameters), finding distinct abductive reasoning capabilities.
TRIDENT proposes a new benchmark to evaluate LLM safety specifically for high-risk domains like finance, medicine, and law, addressing domain-specific compliance.
Research identifies 'over-prompting' in LLMs, where too many in-context examples diminish performance, challenging few-shot learning assumptions.
Research proposes 'Parallel Tokenizers' to improve cross-lingual transfer in encoder models, particularly for low-resource languages, by aligning semantically equivalent words.
Research finds LLMs can generate highly personalized disinformation across languages and demographics, challenging existing safety safeguards.
Researchers propose CSV-Decode, a novel method using sub-vocabularies and geometric bounds to reduce LLM inference compute, maintaining accuracy.
Research proposes ensembling LLM-induced decision trees for error detection in tabular data, aiming to enhance explainability and robustness.
Research explores fine-tuning LLMs to simulate student responses and implicitly model assessment item psychometric properties like difficulty and discrimination.
Researchers propose using top-k log probabilities, a readily available inference signal, to monitor LLM performance and prioritize interventions.
RM-Distiller proposes using generative LLMs more effectively for reward model distillation, moving beyond simple binary annotation.
Research finds adapter merging in two-stage fine-tuned LLMs can reactivate latent reasoning traces, even causing interference post-alignment.
Research explores if aligning RAG models for faithfulness impacts answer accuracy, addressing 'right-answer-wrong-reason' failures in noisy retrieval.
Research introduces Multi-Task GRPO to improve LLM reasoning reliability across diverse tasks, addressing imbalances in standard multi-task RL optimization.
Researchers find frontier LLMs exhibit systematic failure in predicting the outcomes of computational processes and algorithm executions.
Researchers propose an automated evaluator for Text2SQL models to assess translation accuracy on unseen, unlabeled database schemas.
Academic study demonstrates LLMs internally represent accurate token counts even when generating incorrect numerical outputs.
Researchers propose Speculative Pipeline Decoding (SPD) to hide draft-token latency and reduce LLM inference costs via pipeline parallelism.
Researchers propose mechanism-driven monitors to detect LLM training instability early, preventing wasted compute on failed training runs.
Researchers introduce CodexGraph, a framework bridging LLMs and code repositories using graph databases to improve repository-level tasks.
Researchers introduced a mechanistic approach to machine unlearning, using model-internal circuits to explain why some data resists erasure.
Researchers introduced τ-Rec, a benchmark for evaluating conversational, agentic recommender systems using verifiable rewards rather than LLM judges.
An academic practitioner guide detailing the technical architecture, GPU systems, and orchestration layers required to build agentic AI systems.
The FCA secured a confiscation order against convicted fraudster John Burford, recovering a majority of the £1m defrauded from over 100 investors.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion