Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?
Research suggests repairing all missing modalities in multimodal sentiment analysis is not always optimal; some samples perform better with modality subsets.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research suggests repairing all missing modalities in multimodal sentiment analysis is not always optimal; some samples perform better with modality subsets.
Research proposes "Debate-on-Graph" to improve LLM reasoning by addressing noise and errors in knowledge graphs for QA tasks.
Research finds clinical safety evaluations of LLMs in English do not transfer to other languages (e.g., Hausa), especially for smaller models.
EvolvingWorld is a new framework and benchmark for co-evolving AI agents and world models in interactive literary simulations, focusing on long-horizon character and world progression.
Research identifies emerging biosecurity risks from frontier LLMs in scientific workflows, using a specialized bio-red-teaming model and wet-lab validation.
ESCUCHA is introduced as the first Spanish speech understanding benchmark to evaluate large audio language models (LALMs) across diverse acoustic conditions.
Research evaluates SOTA LLMs for citation function classification, achieving new high benchmarks on the ACL-ARC dataset.
VDAR-Router, a new LLM routing method, uses verbalized query difficulty analysis to dynamically select models based on cost-performance trade-offs.
Research proposes Evidence-Grounded Terminology Adaptation (EGTA) for simultaneous speech translation, focusing on recovering paper-specific terminology.
New research introduces Pancasila-Dilemmas, a dataset of 1,834 questions from Indonesian news to evaluate LLM value alignment with country-specific values.
PPL-Factory is a research paper proposing a task-aware and budget-aware method for selecting training data to fine-tune LLMs, improving efficiency.
Research investigates how alignment tuning affects LLM susceptibility to sycophancy and other cue-induced biases by analyzing hidden states.
VEHBench is a new diagnostic benchmark for evaluating LLM performance across different stages of iterative physical engineering design workflows for vibration energy harvesters.
Research evaluates LLM adaptation strategies for climate disclosure classification across varying document sources (e.g., annual reports, press releases).
Research introduces SpecLA, a speculative decoding method for linear-attention models to improve autoregressive decoding efficiency by verifying multiple tokens.
OpenLanguageModel (OLM) is a new open-source PyTorch library designed for building and pretraining small language models with transparent architecture.
MSCE is a training-free framework enabling LLM agents to convert experience into reusable skills and policies, not just memory retrieval.
Research proposes multi-level context modeling to improve expert selection consistency in Mixture-of-Experts (MoE) models, addressing unstable routing.
Research introduces RIMS, a method for preference optimization in RAG for small-scale LLMs, enhancing robustness against noisy retrieval.
Research finds advanced LLMs predict next word similarly to human brain responses measured by EEG, exploring human-like language processing.
Research proposes Parallel Decoder Transformer, a model-intrinsic architecture for generating multiple document sections concurrently from a single LLM.
A new LLM-based indicator developed by PLOS and DataSeer measures research data reuse in scholarly publications, showing a 43% reuse rate.
Research paper argues current coding benchmarks are misaligned with agentic software engineering due to collapsed scoring and lack of component-level signal.
Research introduces GradAlign, a method to improve LLM reinforcement learning performance by selecting high-quality training problems, addressing RL's non-stationarity.
Research explores moving memory inside the LLM agent's observation-reasoning-act loop, allowing constant read/write to address latency challenges.
OmniAgent is a research proposal for a new omni-modal agent architecture designed for efficient long video understanding through active perception.
FormulaCode introduces a new benchmark to evaluate LLM coding agents' ability to optimize entire codebases, moving beyond synthetic, single-objective tasks.
Theoria introduces a verification architecture for AI system answers, converting solutions into auditable, typed state transitions to bridge formal proof certainty with LLM coverage.
FOI-O proposes a global ontology and verification framework for modeling Freedom of Information (FOI) processes using public records.
Research addresses VLM hardware-model mismatch on NPUs, identifying quantization brittleness and I/O-bound autoregressive attention as key issues.
New research proposes PACS, an Implicit Actor Critic Coupling framework, to improve Reinforcement Learning with Verifiable Rewards (RLVR) for LLMs on reasoning tasks.
Research proposes abductive reasoning for generating executable interaction trajectories to overcome data scarcity in RAG agent development.
Research proposes Finetuning-aligned Sequential Training (FAST) for Sparse Autoencoders (SAEs) to address destructive gradient noise in instruct models.
A new arXiv survey reviews Knowledge-Oriented Retrieval-Augmented Generation (RAG) techniques, enhancing LLM accuracy with external knowledge.
New benchmark, PapersPlease, evaluates LLM decision-making in 3,700 moral dilemmas, focusing on prioritizing human needs based on ERG Theory.
Research on Portugal's open-weight national language model, AMALIA, evaluates its trustworthiness as a 'scientific instrument' for community measurement.
New benchmark, MTEB-PT, introduced for evaluating Portuguese text embedding models across 14 datasets, addressing underrepresentation in current metrics.
Research identifies 'phantom transitions' in LLM fine-tuning where cross-entropy loss decreases but correct token ranking fails to improve.
Researchers propose Octo-planner, an efficient on-device Planner-Action framework that separates planning and action execution for AI agents.
Research finds language models with "brittle memory" can re-emit stale, incorrect information confidently if memory retains conclusions but loses supporting work.
Research introduces Dolphin, a decoder-decoder architecture using a 0.5B parameter model to distill long contexts for a 7B LLM, improving energy efficiency.
Research explores fine-tuning LLMs for in-parameter graph reasoning, aiming to overcome token overhead of text-converted graphs.
Research details MiLMMT-46, a Gemma3-based open LLM adapted for multilingual machine translation across 46 languages through continual pretraining and finetuning.
Research explores context compliance in RAG, specifically whether models follow retrieved evidence versus prior knowledge under conflicting information.
Research explores using LLMs to generate synthetic, taxonomy-targeted errors for quantitative reasoning tasks, aiming to aid tutoring and education.
Research finds model performance improves when deeper transformer layers learn context-free value vectors, challenging standard attention paradigms.
Research explored the internal representations of OpenAI's Whisper encoder using sparse autoencoders, finding diverse linguistic and non-linguistic features.
Research investigates why LLMs exhibit overconfidence in incorrect answers, identifying internal mechanisms that cause inflated verbalized confidence.
ClawBench introduces an evaluation framework with 153 real-world online tasks across 144 platforms to test AI agents' ability to complete web-based workflows.
Research identifies 'sockpuppetting' as an enhanced prefill attack, improving LLM jailbreaking through ensembled prefill variants.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion