Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks
Federated learning on clinical text can leak sensitive data via gradient inversion; tokenizer choice impacts privacy risk.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Federated learning on clinical text can leak sensitive data via gradient inversion; tokenizer choice impacts privacy risk.
Research paper proposes an auditable single-system method to evaluate LLM honesty by using a game engine as ground truth, not model self-assessments.
Research benchmarks six MLLMs on a scientific visualization literacy test, finding current evaluations are chart-centric and lack SciVis understanding evidence.
Research explores instruction tuning and model merging to adapt reasoning language models for domains lacking reliable output verification.
Researchers introduced Just Keep Prompting (JKP), a new multi-turn evaluation framework for Vision-Language Models (VLMs) to test stability under sustained questioning.
Research introduces Nous, a belief-based memory architecture for LLM agents using Bayesian inference and information theory for belief revision and forgetting.
Research introduces SD-MAR, a new method using synthetic data and reinforcement learning to improve multi-image analytical reasoning in VLMs.
Researchers introduced Token Time Continuous Diffusion (TTCD), a new diffusion language model operating in continuous space with per-token timings.
Research challenges the assumption that sophisticated prompting and complex datasets consistently improve LLM performance in MCQA tasks.
The EXACT 2026 competition challenges open-weight 8B parameter models to provide explainable answers and logical reasoning for educational QA.
Research proposes a new post-training method for multimodal document Q&A to improve visual grounding without high inference costs or large datasets.
MemoHarness introduces an adaptive agent harness design that learns from experience to improve LLM agent behavior and orchestration.
Researchers propose LLM-T1D, an interpretable language model for closed-loop Type 1 Diabetes control, addressing trust issues in black-box RL systems.
Research introduces a new method for cross-version differencing of scientific documents, addressing challenges with heterogeneous elements and layout.
MedFailBench, a new clinician-built open-source benchmark, evaluates medical AI safety by categorizing errors into specific failure types and severity levels.
Research indicates LLM reliability has an information-theoretic ceiling, meaning perfect reliability is unachievable for any generative task.
Research proposes a statistical self-consistency method for LLMs, aiming to improve reliability by enforcing probabilistic identities like the law of total probability.
Research describes 'harness engineering' – deterministic scaffolding around LLMs for reliable deployment in domain decision systems.
Research paper explores methods for refining LLM-generated research ideas to improve diversity, evaluability, and project execution success.
Research shows fine-tuning LLMs on 'gold answer-conditioned' chains of thought, a common distillation method, degrades verifiable reasoning quality.
Research proposes MARS, a scalable framework for combining LLMs with knowledge graphs (KGs) using multi-hop retrieval and SPARQL generation for grounded answers.
New research introduces CoEvoT, a method for Co-Evolving Chain-of-Thought prompting to improve graph-LLM reasoning, especially under distribution shifts.
Research introduces 'implicit reasoning steering' to bias LLMs towards specific answers without explicit instructions by chaining concepts.
New research introduces PReM, a context compression technique for LLMs that dynamically preserves and refreshes useful information during generation.
DS@GT ARC's LongEval submission evaluates RAG QA systems for citation integrity, using CRAG and CiteFix to correct for divergence in traditional metrics.
CityLLM is a new research framework for natural-language querying of semantic 3D city models, combining spatial and graph databases with LLMs.
Researchers developed T5-CSBoost, an extension of T5-Sentinel, to improve adversarial perturbation resistance in AI-generated text fingerprinting.
A study on arXiv evaluates LLM-generated written corrective feedback (WCF) across 20,000+ EFL essays, emphasizing extrinsic evaluation for learning fit.
Research finds 'structural priors' (cheatsheets) significantly boost in-domain LLM performance for tasks like code security, but degrade out-of-distribution.
Research trains small LLMs to report on internal activation perturbations (activation steering) to detect injected 'thoughts'.
Research proposes a linear-time language identification classifier using compositional data analysis, improving efficiency over neural models.
Research introduces 'Gold-Guided Programmatic Distillation' to improve LLM accuracy in financial reasoning over hybrid tabular and text data.
Research proposes using reasoning graphs to attribute LLM authorship, moving beyond surface-level linguistic features to improve detection robustness.
Research proposes Multi-Head Latent Control, a unified interface for LLM agents to make decisions like deferring, requesting info, or invoking tools.
OmniaBench introduces a new benchmark for evaluating general AI agents across diverse scenarios, tools, and interaction formats.
Researchers at SemEval-2026 Task 11 are addressing LLM limitations in formal reasoning by disentangling logic from content using synthetic data.
Research introduces Expanded Hyper-Connections (xHC), extending Transformer residual streams beyond current limits to enhance memory scaling.
MedRealMM introduces a real-world multimodal benchmark for evaluating LLMs in Chinese online medical consultation, using actual patient data.
Research explores multi-agent LLM framework simulating political coalition formation, addressing RLHF biases that prevent steadfast partisan behavior.
Research proposes SEED, a method for self-evolving on-policy distillation to improve agentic reinforcement learning for LLMs in long-horizon tasks.
Research benchmarks generative AI against supervised Extreme Multi-Label Classification (XMLC) for automated subject indexing of German scientific literature.
Research paper proposes D-cut, an adaptive method for speculative decoding to reduce LLM inference costs under high concurrency by pruning verification depth.
Research introduces "delta signal" for on-policy distillation, an alternative post-training method in reinforcement learning for token-level supervision.
Research audits LLM-generated encyclopedia Grokipedia for political neutrality, comparing it to Wikipedia amidst claims of bias.
Researchers introduced Transformers with Temporal Middle-Layer Recurrence (T2MLR), a new architecture addressing limitations in Transformer reasoning.
Research proposes expanding LLM tokenizers in-place for pre-trained models to improve efficiency and reduce latency for underrepresented languages.
Research introduces ReOPD, an off-environment method for multi-turn on-policy distillation, reducing cost of LLM agent training.
Researchers introduced TikStance, a multimodal dataset of 161 TikTok political videos and 13,876 comments for stance detection.
Researchers propose KV-cache grafting, allowing frozen small models to restore verified knowledge byte-exact, improving capability and reducing cost.
Research introduces Beaver, an agent harness designed for structured scientific curation from multimodal sources including text, tables, and figures.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion