Best-of-Evidence: Best-of-N Selection under Partial Verification
Research proposes "Best-of-Evidence" for model output selection under partial verification, improving on Best-of-N for vision-language tasks.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research proposes "Best-of-Evidence" for model output selection under partial verification, improving on Best-of-N for vision-language tasks.
Research explores latent reasoning in LLMs, specifically whether internal 'scratchpad' assumptions hold under reinforcement learning in complex tasks like chess.
Researchers introduced ADABORD, an AdaBoost framework specifically designed for ordinal classification problems, aiming to better leverage class order.
New research identifies "weight-norm criticality" as a mechanism causing loss spikes during deep neural network training, distinct from learning-rate criticality.
Researchers propose a counterfactual explainability framework using CycleGAN for retinal disease classification to link model decisions to retinal regions.
Research validates using a taxonomy-aware penalty as a training signal for CWE-level vulnerability prediction in Python, improving results.
Research proposes a spectral transformation method to improve LoRA aggregation in federated fine-tuning of Vision Transformers, addressing cross-term errors.
Research explores self-output fine-tuning for autoregressive weather prediction to mitigate error growth over long forecasting horizons.
Research introduces CASC, a new deep subspace clustering method for multivariate spatiotemporal data, focusing on causal dependencies and dynamics.
New research proposes a Reinforcement Learning method to directly optimize the faithfulness of LLM self-explanations, improving how accurately reasoning reflects internal decisions.
New research proposes APEX, a polynomial framework for exact Aumann-Shapley attribution in Graph Neural Networks (GNNs), addressing GNN explainability challenges.
Research proposes using B-Splines to improve the smoothness and training of Neural Temporal Point Processes (TPPs) for modeling event sequences.
Researchers introduced TOUR, a benchmark to evaluate trajectory-level unlearning in Offline Reinforcement Learning (RL) agents, addressing deletion challenges.
New reinforcement learning research proposes Relative Value Learning (RV) to directly learn value differences via an antisymmetric function, rather than absolute values.
Researchers propose Risk Graph Neural Networks (RGNNs) to improve heat-mortality risk estimation by incorporating demographic and geographic context.
New research proposes GKR-HND protocols to verify outsourced Transformer inference, addressing model substitution and incomplete execution risks.
Research presents ARA, an AI framework for automated causal research that identifies invalid causal assumptions to prevent 'silent failures' in analysis.
New research proposes a framework for subgraph filter learning (SFL) to approximate graph filters using partial graph observations, formulated as a statistical learning problem.
Research shows dense prediction rewards cause LLM agents trained with GRPO to fail, driving them to degenerate states rather than task success.
Research explores multi-task learning (MTL) for heterogeneous prediction from video game state data, aiming to improve generalization and reduce costs.
Research finds AI assistants, when used as tutors, can over-intervene, hindering user learning and cognitive engagement during problem-solving.
Researchers developed M$^3$-Gen, a framework to generate interpretable gene expression profiles from clinical and imaging data, addressing data cost and privacy.
Research quantifies the memorization capacity of LoRA adapters, finding it significantly smaller and less predictable than full fine-tuning.
New research suggests model unlearning effectiveness is driven by gradient concentration, not traditional saliency-based weight selection.
Research shows fine-tuning on narrow bad advice can cause broad LLM misalignment by recruiting pre-existing persona structures.
Researchers propose Hilbert Operator for Progressive Encoding (HOPE), a mathematical framework to deconstruct learned representations in deep networks.
Researchers propose "JITAI-Twins," digital twins of subpopulations using diffusion models for mobile health intervention algorithm testing.
Research paper proposes Context-weighted Discrete Flow Matching to improve generative modeling on discrete structures by addressing varying token difficulty.
Research proposes KroQuant, a method for efficient post-training quantization of diffusion transformers to W4A4 without output quality degradation.
Research introduces Test-Time Scaling with error localization to improve LLM inference efficiency by avoiding discarding valid reasoning prefixes.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion