Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
A research survey on continual learning for vision-language models (VLMs) addresses catastrophic forgetting in adapting to non-stationary data.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
A research survey on continual learning for vision-language models (VLMs) addresses catastrophic forgetting in adapting to non-stationary data.
New research proposes a theoretical framework for variance-aware baselines and adaptive learning rates to improve reinforcement learning with verifiable rewards (RLVR) in LLMs.
Research explores continuous-time reinforcement learning for optimal switching problems across multiple regimes, using an exploratory formulation.
Research introduces procedural fairness, defined as equal voice and representation within a policy, for multi-agent multi-armed bandits.
Research proposes a method for continual learning in vision-language models to prevent catastrophic forgetting by preserving semantic geometry.
Open-vocabulary BEV segmentation uses vision-language models for autonomous driving perception beyond training sets, improving real-world robustness.
Research proposes SpecPrefetch, a parameter-efficient expert prefetching method for sparse Mixture-of-Experts (MoE) foundation models to alleviate deployment memory bottlenecks.
Research investigates deep generative models' ability to reproduce complex, non-stationary spatial and spatio-temporal data distributions, a key challenge for real-world application.
Researchers introduce LayerRAG-Bench, a benchmark testing enterprise agentic RAG systems across layer-specific failures like authorization and schema errors.
Research identifies 'Narrative Anchoring' where LLM outputs diverge based on sociolinguistic register despite identical underlying facts.
A new research paper benchmarks LLM performance on logical inference tasks involving probability and uncertainty operators in language.
Researchers evaluated 41 open-weight language models (135M-9B parameters) to benchmark zero-shot intent classification performance and latency.
Researchers introduce B1ade, a minimalist RAG architecture pairing a 335M embedding model with a 1B parameter small language model.
Researchers developed AWARE-FX, a hybrid NLP system combining deterministic logic and financial encoders to extract auditable corporate FX hedging disclosures.
Researchers introduce ReTopK, a method to reduce the computational cost of sparse attention in long-context LLMs by reusing top-K selections.
Researchers propose CoRA, a gradient-free framework converting frozen encoders into task-conditioned retrievers for resource-constrained on-device learning.
Researchers introduce ChronoMem, a framework enabling version control and semantic rollback for LLM agent long-term memory systems.
Researchers propose a framework to ensemble LLM reasoning paths using weighted Directed Acyclic Graphs to improve explainability.
Researchers scaled Memory Decoder to 6.9B parameters, decoupling long-term memory from model reasoning to improve retrieval and storage.
Researchers proposed GGC, a selective query correction framework to improve the semantic accuracy of LLM-generated SPARQL queries over knowledge graphs.
An academic study compares human linguists and LLMs on annotating evaluative language, finding shared difficulties with complex context.
An academic audit reveals that standard LLM-as-a-judge helpfulness rubrics fail to distinguish helpful answers from poor pedagogical guidance.
Researchers developed ParliamentBench, an open-source framework using a social deduction game to evaluate deceptive capabilities in LLM agents.
A research paper demonstrates that compressed language models pass standard fidelity tests but invent invalid steps during agentic execution.
Researchers propose CoMem, a method that caches intermediate-layer states instead of full KV caches, enabling cheaper long-context processing.
Researchers propose a novel evaluation framework to automate LLM output quality assessment without relying on explicit reference standards.
Researchers introduce CACHE-UK, a memory editing method designed to update facts in 4-bit quantized LLMs without degrading performance.
Researchers introduced Fairness Pruning, a lightweight method to locate and mitigate demographic bias in LLM GLU-MLP layers during inference.
Research shows six leading LLMs fail to accurately replicate human belief and stance updates when compared directly to human participants.
Researchers propose using simulated multi-agent persona panels to evaluate Generative UI quality, addressing limitations of single LLM judges.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion