TriQua: Reconciling Granularity and Context in Factuality Evaluation
Researchers introduce TriQua, a factuality evaluation framework that reconciles atomic information extraction with contextual complexity.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduce TriQua, a factuality evaluation framework that reconciles atomic information extraction with contextual complexity.
Researchers introduce QEvict, a method that allows evicted Key-Value cache tokens to be recovered during long-context LLM decoding.
Researchers propose EvoHarness-RL, a reinforcement learning framework to help long-horizon LLM agents manage state and external tools.
Researchers propose a method to automatically learn context-free grammars to guarantee syntactically valid model outputs in domain languages.
Researchers introduce EcoAgent-Bench, a benchmarking framework evaluating LLM agents on their ability to make cost-effective routing decisions.
Researchers propose a verifier-free breadth-depth refinement framework for LLM test-time scaling, reducing reliance on external reward models.
Researchers propose DreamGuard, a runtime guardrail for LLM agents using a world model to simulate and assess multi-step risk before action execution.
Researchers propose Unified Agent, an architecture designed to manage AI agent state and interactions across multiple user devices over time.
Researchers propose predicting task difficulty from text descriptions to optimize agent evaluations without costly simulations.
Researchers identify a performance degradation phase transition in self-evolving LLM agents, where defective distilled skills pollute the agent database.
Researchers propose a framework to refine deep research agent queries by grounding user preferences in structured knowledge graphs.
An academic paper proposes "agentic posture" to address persistent security and authorization gaps in multi-step AI coding agents.
Researchers introduce LangChoiceBench, a benchmark measuring LLM programming language selection, bias, and consistency in code generation.
Researchers introduce HarnessOpt-Bench, a benchmark evaluating LLMs on their ability to optimize agent prompts, tools, and orchestration code.
Researchers introduced AV-AIVAT, a statistical framework reducing agent evaluation costs by up to 74x using anytime-valid stopping.
Researchers propose a pipeline using synthetic data and query decontextualization to improve open-retrieval conversational question answering.
Researchers introduce LELA, a zero-shot, LLM-based entity linking method that bypasses fine-tuning for domain-specific knowledge bases.
Researchers introduce TaxoBench to evaluate deep research agents on their ability to retrieve and organize expert-level taxonomies.
Researchers propose STATe-of-Thoughts to improve diversity and interpretability in inference-time compute using structured action templates.
A research study analyzes over 1,000 court filings containing fabricated AI citations, demonstrating that LLM hallucination rates in legal contexts remain a persistent risk.
Researchers evaluate whether internal-state LLM hallucination detection signals generalize across different languages and domains.
A study of 14 reasoning models reveals that test-time compute scaling fails to reduce factual hallucinations in knowledge-intensive tasks.
Jane Street leads a $2 billion investment in Australian data center operator Firmus Technologies, securing regional AI compute capacity.
Goldman Sachs Research projects global AI investment will exceed $1 trillion by 2026, driven by adjusted hyperscaler capex forecasts.
NatWest Group deployed an AI platform named Serene to detect early signs of financial vulnerability and customer distress.
Apple ML Research published a paper comparing the performance, latency, and arithmetic intensity of diffusion versus autoregressive language models.
Apple ML Research introduced Arbitrage, a speculative decoding technique that reduces reasoning model inference latency and computational cost.
Apple ML Research demonstrates scaling categorical flow matching for discrete data, offering an alternative to autoregressive language models.
LangChain released implementation guidance for establishing user-identity-linked authentication and authorization boundaries for active AI agents.
LangChain published a framework for evaluating agentic AI, covering dataset construction, grader design, and production readiness.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion