Contrastive Conformal Sets
Research introduces Contrastive Conformal Sets, extending conformal prediction to contrastive learning for distribution-free coverage guarantees in feature spaces.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces Contrastive Conformal Sets, extending conformal prediction to contrastive learning for distribution-free coverage guarantees in feature spaces.
Research on preference encoding in looped transformers found original headline results for evaluator accuracy were inflated due to evaluation errors.
Research identifies domain shift as a major challenge for machine learning-based Android malware detectors in real-world deployments.
Researchers introduced Grow-Prune-Freeze (GPF) networks, an adaptive continual learning technique for dynamic, non-stationary tasks like olfactory navigation.
Methodological study analyzes collinearity, dimensionality, and cluster stability using PCA and k-means clustering on US airline profit cycles.
Research finds code correctness is linearly decodable from Qwen3-4B-Instruct-2507 LLM hidden states before generation, using LiveCodeBench.
Research proposes zero-shot quantization for object detection models using off-the-shelf generative models, addressing data access limitations.
Researchers propose Cross-Cluster Weighted Forest (CCWF), an ensemble method addressing data heterogeneity for improved accuracy and generalizability.
Research explores distributionally robust optimization (DRO) for robust inference in continuous probability spaces using iterative algorithms.
Research introduces a stochastic smoothing framework for nonconvex-nonconcave minEmax problems, relevant to Wasserstein distributionally robust optimization.
Research investigates methods to avoid and reverse model collapse when retraining generative models on synthetic data through verification.
Research introduces DualHNIE, a hypergraph learning method for estimating node importance in heterogeneous knowledge graphs by capturing higher-order interactions.
Research explores neural network architectures for amortized Bayesian inference, leveraging deep learning for statistical modeling.
Research finds that motion imitation learning (IL) alone struggles to generate biomechanically consistent joint moments without kinetic data.
Research identifies AgentWorm, a self-propagating attack vector across LLM agent ecosystems like OpenClaw, exploiting tool execution and messaging.
Research explores unsupervised evaluation of deep audio embeddings for music structure analysis, mitigating reliance on annotated data.
Research proposes learning multi-vulnerability attack chains in software supply chains from SBOM graphs, moving beyond per-CVE analysis.
Research paper SLYP demonstrates LLM agents can reason about vulnerabilities in commercial off-the-shelf (COTS) binary software.
Research paper introduces Deep Nash Q-Network (DNQ) for solving partially observable n-player games, tested on multi-turn simultaneous bidding.
Conan-embedding-v3 proposes a decouple-fuse-recover framework to create unified omni-modal embeddings from text, image, video, and audio.
arXiv paper proposes a framework for generative AI risk control in financial institutions, moving beyond conversational AI to specialized applications.
Research analyzed 25,264 agentic pull requests across 2,361 GitHub projects to understand early adoption and management of agentic coding tools.
Researchers introduced Just Keep Prompting (JKP), a new multi-turn evaluation framework for Vision-Language Models (VLMs) to test stability under sustained questioning.
New research introduces LBA, a method for generating high-quality adversarial texts with significantly lower query budgets in hard-label scenarios.
Research explores latent communication channels between LLM agents, suggesting complex concepts exceed text-based expressibility, with implications for multi-agent systems.
Research explores automatically evolving task-specific prompt guidelines to address underspecified user queries and improve LLM reliability.
Researchers introduced Token Time Continuous Diffusion (TTCD), a new diffusion language model operating in continuous space with per-token timings.
Research introduces Polestar, a method for improving diffusion LLM inference efficiency by addressing bidirectional attention and KV-cache reuse issues.
Research introduces "tool efficiency" and "marginal tool utility" as new quantitative metrics to evaluate the rate and usefulness of LLM agent tool calls.
Research challenges the assumption that sophisticated prompting and complex datasets consistently improve LLM performance in MCQA tasks.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion