Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding
Researchers propose Chopthin-Consensus Power Sampling, an inference-time decoding method to improve LLM reasoning and maintain trajectory diversity.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers propose Chopthin-Consensus Power Sampling, an inference-time decoding method to improve LLM reasoning and maintain trajectory diversity.
Researchers propose EAR, an Entity-Aware Partitioning method for RAG that replaces fixed-size chunking with entity-anchored units.
Researchers propose CueMem, a cue-guided framework for long-term conversational memory that uses extracted records as retrieval cues.
Researchers propose ORQA, a framework for evaluating LLM professional knowledge using O*NET occupations and trusted regulatory websites.
ArXiv paper introduces GraphProfiler, using LLMs and personal knowledge graphs to infer sensitive attributes from indirect post data.
Researchers propose an adaptive inference layer to reduce false wake-up activations in conversational AI systems using post-ASR correction.
Researchers introduced AMDKernelVault, an open HIP and Triton kernel corpus and training framework for AMD CDNA GPUs.
Researchers introduced Zipbench, a framework designed to compress comprehensive LLM benchmark suites to reduce evaluation costs.
Researchers propose LifeMem, a lifelong learning framework designed to help LLM agents reuse experience and mitigate catastrophic forgetting.
Researchers propose Cognition on Graph, a framework combining cognitive cycles and graph-text synergy to improve RAG reasoning.
Researchers propose a three-stage training pipeline for compact, efficient dense retrievers supporting Polish and European languages.
Research shows binary-choice LLM benchmarks like TruthfulQA suffer from surface-level feature leakage that inflates accuracy scores.
Research reveals LLMs struggle with long-horizon procedural reasoning when executing complex tasks using multi-page manuals.
Researchers propose Duplex Cue, an evaluation framework for in-turn adaptation in full-duplex voice agents handling overlapping speech.
Researchers propose PRISMA-LLM, a reporting framework to audit LLM and software workflows in systematic reviews, analyzing 888 papers.
Research evaluates agentic coding harnesses using a private, contamination-controlled suite to isolate harness efficacy from model capabilities.
A benchmark study compares general shell interfaces against typed tools for enterprise digital worker agents across business tasks.
Researchers examine how draft-verify-revise LLM pipelines handle deictic ambiguity when context cascades between multiple models.
Researchers introduced LAST-CQ, an agentic Text-to-Cypher framework, testing counterfactuals to measure which loop components drive query recovery.
Researchers propose MAxBench to evaluate multinomial concept steering and representation geometries in language models beyond binary tasks.
Researchers demonstrate that current multilingual LLM watermarking methods fail robustness tests under translation attacks in low-resource languages.
Research demonstrates spurious training data correlations drive LLM hallucinations and impair automated detection tools.
Researchers introduce MMGR, a benchmark to test whether multimodal generative models perform genuine physical and logical reasoning.
Researchers propose PACIFIC, a framework using psychometric personality traits to improve preference alignment in large language models.
Research analyzes pragmatic framing in LLM prompts, demonstrating how authority or urgency cues alter outputs without changing tasks.
Researchers introduce ReasoningFlow, a framework using directed acyclic graphs to map non-linear LLM reasoning traces like self-correction.
Researchers propose SAEExplainer, a framework using mechanistic feedback to improve sparse autoencoder feature interpretation in large language models.
Researchers identify a retrieval-conditioned rebinding circuit in LLM attention heads that manages dynamic state tracking and entity attributes.
A systematic study on arXiv investigates the trade-offs between effectiveness and text fluency when conditioning LLMs for concept control.
Researchers introduced GRACE-DS, an isolated evaluation framework for pre-deployment testing of LLM-powered AutoML agents.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion