Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
Researchers propose a zero-shot LLM detection method via latent prompt restoration to identify machine-generated text.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose a zero-shot LLM detection method via latent prompt restoration to identify machine-generated text.
An arXiv research paper evaluates speech LLMs against context biasing methods for recognizing rare and domain-specific words in speech recognition.
Researchers introduce HallDetect, a reference-free, black-box framework for detecting LLM hallucinations in source-grounded tasks.
Academic researchers introduced Synthetic Query Probing, a method to map similarity score distributions across different embedding models.
Researchers propose MACRO, a Markov chain routing method for transformer layers to dynamically skip or repeat layers without retraining.
Researchers prove that routing web agent tasks to the optimal visual/text observation mode is mathematically difficult to learn.
Researchers benchmark LLM capabilities in reviewing complex, rule-intensive national standard documents to improve consistency and compliance.
Researchers introduced a reference-free framework using LLM judges to evaluate the quality, consistency, and complexity of conversational agent benchmarks.
Academic research evaluates programmatic code execution against structured JSON tool-calling, highlighting trade-offs in agent capabilities.
Researchers introduce MIST, a benchmark evaluating an LLM's ability to selectively trust external context over internal knowledge.
An academic paper introduces Agentic Nesting, a methodology to orchestrate and integrate fragmented legacy enterprise applications using LLMs.
Researchers introduce TriQua, a factuality evaluation framework that reconciles atomic information extraction with contextual complexity.
Researchers introduce QEvict, a method that allows evicted Key-Value cache tokens to be recovered during long-context LLM decoding.
Researchers propose EvoHarness-RL, a reinforcement learning framework to help long-horizon LLM agents manage state and external tools.
Researchers propose a method to automatically learn context-free grammars to guarantee syntactically valid model outputs in domain languages.
Researchers introduce EcoAgent-Bench, a benchmarking framework evaluating LLM agents on their ability to make cost-effective routing decisions.
Researchers propose a verifier-free breadth-depth refinement framework for LLM test-time scaling, reducing reliance on external reward models.
Researchers propose DreamGuard, a runtime guardrail for LLM agents using a world model to simulate and assess multi-step risk before action execution.
Researchers propose Unified Agent, an architecture designed to manage AI agent state and interactions across multiple user devices over time.
Researchers propose predicting task difficulty from text descriptions to optimize agent evaluations without costly simulations.
Researchers identify a performance degradation phase transition in self-evolving LLM agents, where defective distilled skills pollute the agent database.
Researchers propose a framework to refine deep research agent queries by grounding user preferences in structured knowledge graphs.
An academic paper proposes "agentic posture" to address persistent security and authorization gaps in multi-step AI coding agents.
Researchers introduce LangChoiceBench, a benchmark measuring LLM programming language selection, bias, and consistency in code generation.
Researchers introduce HarnessOpt-Bench, a benchmark evaluating LLMs on their ability to optimize agent prompts, tools, and orchestration code.
Researchers introduced AV-AIVAT, a statistical framework reducing agent evaluation costs by up to 74x using anytime-valid stopping.
Researchers propose a pipeline using synthetic data and query decontextualization to improve open-retrieval conversational question answering.
Researchers introduce LELA, a zero-shot, LLM-based entity linking method that bypasses fine-tuning for domain-specific knowledge bases.
Researchers introduce TaxoBench to evaluate deep research agents on their ability to retrieve and organize expert-level taxonomies.
Researchers propose STATe-of-Thoughts to improve diversity and interpretability in inference-time compute using structured action templates.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion