Convolution for Large Language Models
Research explores using depthwise convolutions within Transformer blocks to improve local inductive bias in LLMs, comparing placements across Qwen3.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research explores using depthwise convolutions within Transformer blocks to improve local inductive bias in LLMs, comparing placements across Qwen3.
A collaboration between the DGT and EMT Network is localizing the MMLU dataset into 11 European languages to create a multilingual LLM evaluation benchmark.
Relay-Bench introduces a new benchmark for evaluating LLMs on multi-domain reasoning chains, with GPT-5.5 (xHigh) scoring 43.3%.
New research proposes ScAffolded Generative models for Explanation (SAGE) framework for computational pragmatic reasoning in language models.
Research explores adapting LLM-based classification, developed on US police data, to identify vulnerability indicators in UK police incident logs.
Research finds requesting JSON output from 44 LLMs significantly reduces answer diversity on open-ended prompts, collapsing options.
Search-on-Graph-R1 trains a compact 8B LLM to navigate knowledge graphs for question answering using SFT and RL, aiming to reduce inference costs.
Research finds reasoning fine-tuning induces persistent, global internal changes in LLMs for multi-step reasoning, not just local token competence.
Research finds narrative framing significantly influences LLM agent behavior, often more than explicit persona prompts, across structurally identical tasks.
Research from arXiv investigates how instruction-tuned Transformer models, LLaMA and Mistral, encode causation and antithesis.
Research explores machine unlearning for vision-language models (VLMs), noting that language backbone unlearning doesn't guarantee VLM unlearning due to visual data influence.
LatentMT explores latent-reasoning LoopLMs for machine translation, adapting a small 2.6B-parameter model with recurrent computation.
Fusion Embedding proposes a unified embedding space for text, images, video, and audio, integrating audio into existing vision-language models.
Research explores rationale-guided knowledge distillation for cross-lingual stance detection, improving model performance in low-resource languages.
Research introduces FiT, a diagnostic method to select small LLMs for fine-tuning in critical QA tasks like cybersecurity, before costly adaptation.
RF-Agent is a new framework using textbook-driven knowledge distillation and a multi-agent QTSA pipeline for RFIC design, addressing data scarcity.
New research proposes Causal Alignment and Structural Enforcement (CASE) to improve chain-of-thought faithfulness in LLMs, ensuring reasoning supports answers.
Research introduces AILQA, an AI system for Indian legal question answering, leveraging LLMs to navigate complex Indian legal texts.
Research proposes Hierarchical Parallel Document Parsing (HPD-Parsing) to address sequential bottlenecks in unified VLM-based document parsers.
NVIDIA Nemotron 3.5 ASR 0.6B was adapted for Kikuyu, Dholuo, and Kalenjin languages, demonstrating data-centric methods for low-resource ASR.
New research proposes Step-Level Self-Consistency Group Relative Policy Optimization to reduce hallucinations in multi-step LLM reasoning.
Research shows ASR models encode both verbatim and intended transcription styles, but uncontrolled activation causes decoding instability and unreliable word timing.
Research explores self-evolution of LLM dialogue skills using future-feedback prediction to address unstable validation signals in open-ended conversations.
AutoJourn is a research system demonstrating multi-perspective summarization, bias detection, and bias neutralization for LLM-generated news.
Research proposes a new taxonomy for curriculum learning in NLP, disentangling difficulty evaluation from training scheduling to improve analysis.
Researchers introduced MedDDC-Eval, a diagnosis-decoupled evaluation framework for multi-turn medical consultation agents to isolate policy elicitation from diagnosis generation.
New research proposes PINT, an invariant speech tokenization method that disentangles semantic content from non-linguistic speech variations.
DAIS (Dependency-Aware Intermediate QA Supervision) is a new training framework converting teacher rationales into stage-level QA records to improve complex reasoning.
Research explores using translated data to augment scarce expert-annotated corpora for text difficulty assessment, particularly for lower-resource languages.
Research explores enhancing legal machine translation using reasoning-capable language models to improve precision and address linguistic complexity.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion