CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
Research identifies a benchmark, CiteAudit, to detect hallucinated citations from LLMs, which are present in scientific submissions.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research identifies a benchmark, CiteAudit, to detect hallucinated citations from LLMs, which are present in scientific submissions.
Research challenges the 'ground truth' paradigm in data annotation, arguing human disagreement is a critical signal, not noise, for ML model training.
Research paper claims current LLMs fail to grasp user intent beyond explicit harmful content, creating exploitable security vulnerabilities.
Research finds fine-tuning LLMs on synthetic data from diverse sources mitigates distribution collapse, adversarial robustness, and self-preference bias.
Research introduces AgentSeer, an observability tool decomposing agentic executions into action-component graphs to quantify model-level and agent-level risk gaps.
Research indicates agent harnesses, not just the LLM, contribute significantly to an agent's competence and performance.
Research introduces JudgeSense, a benchmark measuring LLM-as-a-judge system sensitivity to semantically equivalent prompt paraphrases, via a Judge Sensitivity Score (JSS).
Research explores using distillation and reinforcement learning to enable compact LLMs (0.5-1B parameters) to perform agentic RAG and search behaviors.
AgentHER improves LLM agent performance by relabeling failed trajectories as successful for different goals, recovering lost training data.
Research identifies fragmented benchmarks for Large Audio-Language Models (LALMs) and proposes a systematic taxonomy for comprehensive evaluation.
A research audit of 1.5K American newspapers found 9% of articles published in summer 2025 were partially or fully AI-generated, with uneven disclosure.
Research systematically benchmarks context utilization techniques (CMTs) for language models, addressing issues of ignored or irrelevant information.
Research evaluated agentic LLMs on synthesizing longitudinal multiple myeloma patient records against expert clinical consensus for treatment decisions.
Researchers propose a method for dynamically routing LLM queries to specific attention heads for re-ranking, improving relevance estimation.
Research identifies adversarial instruction vulnerabilities in LLM applications like resume screening; defenses for specialized domains lag behind core areas.
Research demonstrates a new 'intention deception' method for jailbreaking frontier LLMs, exploiting brittleness in current safety alignment.
Research proposes an evaluation framework for highlight explanations, aimed at showing which context pieces LMs use to generate responses.
Research proposes using lightweight probes on LLM hidden states to perform classification tasks like safety filtering within the same forward pass.
AgentPulse introduces a continuous evaluation framework for AI agents, scoring 50 agents across 10 categories using 18 real-time deployment signals.
Research tested if humans detect AI-assisted writing and if AI detection warnings influence human writing with chatbots.
RCSB PDB implemented a retrieval-augmented generation (RAG) system for its help desk to assist expert biocurators with protein structure deposition.
Xiaohongshu's RedParrot system improves NL-to-DSL conversion for business analytics using query semantic caching to reduce LLM latency and cost.
Research proposes STELLAR-E, a synthetic data generator for rigorous, domain-specific, and language-specific LLM evaluation, addressing privacy and data scarcity.
AgentEval proposes a DAG-structured framework for evaluating agentic workflows, tracking error propagation at each step to improve reliability.
Research proposes CRISP, a sparse autoencoder method for persistent concept unlearning in LLMs, aiming to remove unwanted knowledge from model parameters.
Research paper proposes an isolation-first, containerized architecture for secure on-premise deployment of open-weight LLMs in radiology.
Research introduces StorySim, a framework generating synthetic stories to evaluate LLM Theory of Mind and world modeling without data contamination.
Research investigates if LLMs track source trustworthiness in Turkish evidential morphology, finding humans show robust trust sensitivity, LLMs less so.
Research proposes "Green Shielding," a user-centric approach to build deployment guidance for LLMs by characterizing how benign input variation shifts model behavior.
Researchers introduced For-Value, a forward-only data valuation framework for LLMs and VLMs, enabling efficient, batch-scalable finetuning.
DepthKV proposes a new KV cache pruning method for LLMs, reducing memory footprint linearly with sequence length, optimizing long-context inference.
Research explores structural pruning techniques to compress existing Large Vision Language Models (LVLMs) for deployment on resource-constrained devices.
Research identifies methods for deliberately aligning LLMs with specific political ideologies through prompt engineering or fine-tuning, raising misuse concerns.
OS-SPEAR is a new research toolkit for evaluating OS agents' safety, performance, efficiency, and robustness, addressing current benchmark limitations.
Research proposes a novel method, "Layerwise Convergence Fingerprints," for real-time detection of LLM misbehavior like jailbreaks and prompt injections.
FinGround is a new research method to detect and ground financial hallucinations in LLMs by verifying atomic claims against regulatory filings, improving accuracy by 43%.
Research quantifies inter-LLM divergence in API discovery and ranking across 15 domains and 5 model families, impacting agent reliability.
K-MetBench introduces a multi-dimensional benchmark for evaluating expert reasoning, locality, and multimodality in LLMs for meteorology.
New research proposes a lightweight method for extracting visual elements from PDFs, including figures, tables, and forms, improving RAG performance.
ShredBench evaluates Multimodal LLMs on document reconstruction from shredded fragments, a challenging task requiring semantic and visual integration.
Research presents VeriLLMed, an interactive visual debugging tool using knowledge graphs to assess medical LLM diagnostic reasoning reliability.
Researchers identify 'conditional system prompt poisoning' (PARASITE) as a supply-chain vulnerability in LLMs, allowing malicious code injection via prompts.
SWE-Pruner proposes a self-adaptive context pruning method for LLM coding agents to reduce API costs and latency by focusing on task-specific code understanding.
Research paper argues that logical soundness is not a reliable criterion for neurosymbolic fact-checking with LLMs, challenging a common mitigation strategy.
Research introduces SpeechLLMs for direct speech processing, questioning if it improves speech-to-text translation quality over cascaded methods.
Research paper SWE-QA introduces a new benchmark for evaluating LLMs' ability to answer complex, repository-level code questions beyond simple snippets.
Research identifies prompt underspecification as a key source of LLM instability, leading to significant performance degradation when prompts or models change.
Researchers introduced an N-gram Coverage Attack, a membership inference method effective against API-only LLMs like GPT-4, without hidden state access.
AdaComp is a new context compression method for RAG that uses an adaptive predictor to extract relevant sentences, aiming to reduce noise and cost.
Research identifies a flaw in audio-language model evaluation: models can achieve high scores on audio benchmarks using text priors, not true audio understanding.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion