Watermark Forensics for Generative Models: An Information-Theoretic Perspective
Research explores advanced watermarking for generative models, moving beyond detection to attribution, payload extraction, and localization within edited text.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research explores advanced watermarking for generative models, moving beyond detection to attribution, payload extraction, and localization within edited text.
A new Bayesian filtering technique improves state estimation by using processor-native uncertainty tracking for inference in sensor-rich systems.
Research proposes adaptive self-distillation during SFT to prevent overconfidence and enhance generation diversity in large reasoning models.
Researchers propose TTARAG, a test-time adaptation method to improve Retrieval-Augmented Generation (RAG) performance in specialized domains by dynamically updating the language model.
UXBench introduces a new user-centric benchmark for AI assistant evaluation, using real user feedback to measure preference and dialogue quality.
Research introduces AI-assistability (AI), a new metric combining structural alignment and functional correctness to benchmark multi-agent frameworks.
Research applies Signal Detection Theory to measure LLM metacognitive efficiency, separating knowledge from confidence signal tracking.
Research unifies observed mechanistic interpretability phenomena in Transformers by proposing hierarchical latent structures in data generation.
Research explores using public sentiment analysis for Advanced Air Mobility (AAM) to identify societal barriers and aid adoption strategies.
NAVER LABS re-implemented its IWSLT 2025 instruction-following pipeline for the 2026 task, adapting to SeamlessM4T-v2-large and Qwen3-4B-Instruct.
Research compares human and LLM (Llama-3.3-70b-versatile, GPT-4o-mini) political ideology annotations, finding topic sentiment affects perception.
Research introduces Sibling-Guided Credit Distillation (SGCD) to improve long-horizon tool-use reinforcement learning by refining credit assignment.
Git-Assistant combines LLMs with formal reasoning for improved Git repository management, targeting common developer challenges.
EntSQL is a new benchmark for text-to-SQL models, specifically designed to evaluate their ability to use private enterprise knowledge in long contexts.
Research highlights that supervised financial NLP benchmarks used for model selection suffer from 'measurement risk,' where rubric, metric, or aggregation choices significantly influence results.
Research explores if LLM 'personality' prompting, affecting communication style, impacts objective task outcomes in multi-agent teams.
Research finds audio-language embedding models like CLAP struggle with negation, encoding affirmative and negated captions similarly.
FAIR GraphRAG proposes a retrieval-augmented generation approach leveraging knowledge graphs to enhance LLM responses to domain-specific questions.
Research proposes Malleable Prompting, an interactive technique converting natural language prompt preferences into GUI widgets for controllable LLM generation.
Research proposes Automated Reasoning checks (ARc) using LLMs to formalize natural language policies for improved correctness guarantees.
Research introduces PalmClaw, an on-device agent framework for mobile phones enabling LLM agents to perform multi-step tasks locally using device tools and data.
Research indicates LLM judges are overly generous in evaluating open-ended responses when no reference answer is provided, impacting reliability.
Research presents the first meta-evaluation of LLMs generating rubrics for assessing open-ended outputs, aiming to scale experiment reproduction.
Research explores co-evolving evaluation metrics and LLM agent skills for self-improving systems, claiming metrics can also be evolved.
Research proposes "Knowledgeless Language Models" (KLLMs) that suppress parametric recall to force evidence-grounded reasoning, aiming for reliability.
Research explores efficient inference techniques for Masked Diffusion Large Language Models (dLLMs) to achieve practical speedups over autoregressive models.
Researchers propose QDEvo, a multi-objective framework integrating LLMs with evolutionary computation to improve heuristic design for combinatorial optimization by addressing mode collapse and enhancing semantic diversity.
Research finds LLMs maintain aggregate accuracy on benchmarks but exhibit significant prediction flips when task-irrelevant context is added.
Research evaluates LLMs' ability to identify and correct patient misconceptions over multi-turn medical conversations, highlighting current evaluation framework gaps.
Research tested 44 LLMs on open-ended single-word prompts, finding significant answer convergence (e.g., "serendipity" 41% of the time).
Research explores cost-aware speculative decoding for Mixture-of-Experts models, aiming for faster inference by optimizing expert activation.
Research explores using Quadratic Unconstrained Binary Optimization (QUBO) with unconventional solvers for evidence selection in RAG, moving beyond top-k methods.
Research integrates Small Language Models (SLMs) with a culturally-sensitive Responsible NLP Framework to detect health misinformation in low-resource languages.
Research compares semantic search dynamics in humans and LLMs (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency and NLP metrics.
Research evaluates agentic LLM systems for breast cancer treatment recommendations across 72 clinical cases using 1,147 rubrics.
Research paper introduces a method to measure LLMs' ability to shift between expressing an expert consensus and their own stance.
Research developed a method to fine-tune Llama 3 (8B) into an efficient cross-encoder for RAG reranking using knowledge distillation and quantization.
Research proposes a text dataset distillation framework, reducing corpora to 0.1% of original size while preserving downstream task fidelity.
Research introduces CARE-PPO, a reinforcement learning framework using PPO fine-tuning for LLMs to estimate confidence in quantitative predictions.
New research proposes SeRIn, a multimodal LM fusion scheme that separates modality-specific refinement from cross-modal integration for sentiment analysis.
Researchers introduced CANDI-QA, a new benchmark to evaluate LLMs for contextual understanding in specialized domains like finance and medicine.
Research explores sparse inter-layer dependencies within Transformer FFNs, introducing a training-free attribution method for interpretability.
Research evaluates citation faithfulness of a 4B parameter model performing on-device research, differentiating between claim faithfulness and coverage.
New research proposes JoLT, a near-lossless KV cache compression method for LLMs, addressing memory limits and improving inference throughput.
New research proposes a principled framework for evaluating extractable memorization in LLMs, distinguishing it from general text reproduction.
Research questions the necessity of large parameter counts (e.g., >1B) in multimodal emotion language models for performance.
Research finds LLM outputs are highly concentrated and narrow compared to human responses, exhibiting a short-tail distribution.
Research finds LLMs' sequential decision-making can be biased by induced emotions, using the Iowa Gambling Task as a testbed.
Research explores translating low-resource languages to English for fine-tuning English BERT models, addressing data and computational challenges.
Research identifies context window limitations and noisy content accumulation as primary barriers to scaling long-horizon agentic search systems.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion