LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
Research introduces dynamic LoRA selection and merging at inference time to adapt large language models to diverse, unpredictable tasks without re-training.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces dynamic LoRA selection and merging at inference time to adapt large language models to diverse, unpredictable tasks without re-training.
Research benchmarks LLM bias in multilingual financial misinformation detection, identifying behavioral biases from human-authored training data.
Matrix is an arXiv research paper proposing a peer-to-peer multi-agent framework for synthetic data generation, removing centralized orchestration.
Research fine-tunes Qwen2.5-3B-Instruct to use Prolog as an external symbolic reasoning tool to improve accuracy and verifiability.
Research compares Agentic RAG and standard RAG, finding Agentic RAG marginally better for complex questions but with higher cost and latency.
Research finds frontier LLMs struggle to generate statistically valid random numbers from specified distributions, failing fundamental probabilistic sampling tests.
Research proposes TWGuard, an approach to optimize LLM safety guardrails for specific linguistic and cultural contexts to improve in-the-wild effectiveness.
New benchmark, BengaliMoralBench, created to audit moral reasoning in LLMs for Bengali language and culture, addressing Western bias.
Research paper introduces SaLAD, a multimodal safety benchmark with 2,013 real-world image-text samples across 10 common scenarios, to evaluate MLLM safety.
Research indicates sparse attention algorithms, intended for LLM inference efficiency in the decode stage, can degrade performance.
A research survey on Table Question Answering (TQA) methods, tasks, and evaluation, noting recent LLM advances and remaining systematic challenges.
Research introduces StealthGraph, a knowledge-graph-guided method to generate domain-specific harmful prompts for LLM red-teaming, focusing on implicit risks.
Research finds benign fine-tuning can cause LLMs to lose contextual privacy reasoning, leaking sensitive data even with subtle training patterns.
Research finds fine-tuned LLM-as-a-judge models degrade over time with new data, impacting future-proofing and backward-compatibility.
ReTraceQA proposes a new benchmark to evaluate reasoning traces, not just final answers, for Small Language Models (SLMs) in commonsense QA.
Research suggests adapting tool schemas to Small Language Models (SLMs) improves tool-use performance in multi-agent systems, reducing hallucination.
New arXiv paper proposes an alignment algorithm to evaluate speech recognition systems, focusing on semantically weighted errors in rare terms and named entities.
Research identifies performance gaps in LLM-based rerankers for cold-start recommender systems, citing coverage and exposure issues.
Research personalizes LLMs to emulate judicial reasoning using synthetic-organic supervision for fine-tuning in low-resource settings (Hebrew).
MegaRAG proposes combining knowledge graphs with RAG to improve LLM high-level conceptual understanding and deep reasoning over long documents.
Research proposes "functional fragmentation" for LLM-as-a-Judge evaluations, breaking outputs into rhetorical functions for granular scoring.
Research finds automated evaluation of LLM agents is unreliable, with errors propagating through tool-use chains. Benchmarked 9 LLMs.
Research finds LLMs struggle with human-like, structure-sensitive world knowledge integration in ambiguity resolution, unlike humans.
New benchmark, Text2DistBench, evaluates LLMs' ability to infer distributional knowledge from text collections, moving beyond single-fact extraction.
Research introduces Reasoning Memory, a retrieval-augmented method improving LLM reasoning by reusing procedural knowledge from prior problem-solving trajectories.
Research finds large vision-language models (LVLMs) and humans use different grounding mechanisms in multi-turn referential communication tasks.
Research identifies positional and language biases in long-document embeddings, impacting discoverability of document segments.
Research evaluates LLM adherence to counterfactual medical evidence vs. model priors, using a new MedCounterFact QA dataset.
Research indicates LLMs internally encode token-level functional importance within reasoning chains, potentially enabling more efficient compact reasoning.
Researchers propose TLoRA, a new LoRA variant that optimizes rank allocation, scaling, and initialization to improve parameter-efficient fine-tuning.
Researchers demonstrated Factorized Linear Projection (FLiP) models can recover over 75% of lexical content from multimodal, multilingual sentence embeddings.
HPLT 3.0 presents an open, 30-trillion-token multilingual dataset for LLM pre-training, covering almost 200 languages.
Research finds emergent misalignment (EM) can occur in LLMs via in-context learning, not just finetuning, across Gemini, Kimi-K2, Grok, and Qwen.
Research indicates LLMs may use 'choices-only' strategies in multiple-choice questions, even with reasoning steps, raising concerns about true understanding.
Research critiques medical diagnostic LLM benchmarks, citing contamination bias from public exams and lack of real-world clinical complexity.
Research finds LLMs, like humans, conflate logical validity with semantic plausibility, revealing a bias in reasoning mechanisms.
Research explores how training data quantity and quality affect LLM arbitration between parametric knowledge and in-context information when they conflict.
Researchers released ToxiFrench, a 53,622-comment dataset for French toxicity detection, benchmarking models via CoT fine-tuning.
Research formalizes "user-assistant bias" in LLMs, where role tag asymmetries in training data introduce inductive biases affecting model behavior.
Research identifies a vulnerability where a single user can persistently alter LLM knowledge via selective upvoting/downvoting of stochastic model outputs.
Researchers applied clinical personality assessment validity scales (L, K, F, Fp, RBS) to 20 frontier LLMs' metacognitive self-reports across 524 items.
Research proposes using data compressibility to quantify LLM memorization, offering a new method to measure training data influence.
Research paper introduces LTRR, a learning-to-rank framework for dynamically selecting optimal retrievers in RAG systems based on query type.
Research finds frontier LLMs excel at lexical code recall but struggle with semantic understanding and operational semantics in long code contexts.
LLMs show enhanced robustness against individual simple biases but remain vulnerable to ensembles of multiple biases in real-world data, leading to unstable performance.
Research indicates LLMs maintain consistent value orientations despite persona prompting, showing inertia in moral and value judgments.
Research proposes uncertainty-calibrated fine-tuning to reduce LLM hallucinations and improve reliability by estimating response confidence.
Research finds multi-agent LLM systems for open-ended idea generation exhibit 'diversity collapse' due to structural coupling, limiting solution space.
Research identifies sparse autoencoder (SAE) features in LLMs that reveal semantically coherent, context-consistent network components.
DuConTE, a new dual-granularity text encoder with topology-constrained attention, improves text-attributed graph processing over existing LM/GNN methods.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion