AutoNFS: Automatic Neural Feature Selection
AutoNFS proposes a neural feature selection method that automatically determines the optimal number of features for tabular data without user intervention or retraining.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
AutoNFS proposes a neural feature selection method that automatically determines the optimal number of features for tabular data without user intervention or retraining.
QuantSightBench evaluates LLMs on quantitative forecasting tasks with prediction intervals, moving beyond simple judgmental questions.
Research claims mixed precision settings in distributed deep learning can cause training time variations of ~2.4x; existing prediction models lack this capture.
Research paper reviews State Space Models (SSMs), including Mamba, highlighting their linear scaling, long-range dependency capabilities, and efficiency.
Research presents evidence for hybrid recurrent-attention neural networks outperforming pure transformers, specifically the Olmo Hybrid model.
ConFu, a new speculative sampling method, uses a multi-branch predictor to improve draft model quality, enhancing LLM inference speed.
Research introduces TabularMath, a benchmark for evaluating LLMs on multi-step mathematical reasoning over tables, including incomplete data.
Research evaluates large language model robustness to errors in Chain-of-Thought reasoning steps, finding specific perturbation types degrade performance.
RedBench is a new universal dataset for red teaming large language models, aggregating 37 existing benchmarks for systematic vulnerability assessment.
Research claims supervised fine-tuning (SFT) can increase LLM hallucinations due to new factual exposure, proposing continual learning to mitigate this.
Research proposes an open-ended Arabic cultural QA benchmark with dialect variants, converting MCQs to OEQs to evaluate LLM performance.
OjaKV introduces context-aware online low-rank compression to reduce KV cache memory usage for long-context LLMs, addressing a significant inference bottleneck.
TRIDENT proposes a new red-teaming dataset synthesis method for LLM safety, focusing on tri-dimensional diversity beyond lexical variation.
Research investigates the disconnect between interpretability and semantic correctness in Chain-of-Thought (CoT) traces used in LLM knowledge distillation.
Research uses perturbation-based attribution to compare interpretive behaviors of LLMs for automated code compliance across fine-tuning strategies.
Research indicates Vision-Language Models (VLMs) may primarily leverage text reasoning over true vision-grounded reasoning, impacting multimodal task reliability.
Research introduces DELEGATE-52 benchmark to assess LLMs' ability to maintain document integrity in long, delegated workflows, identifying error introduction.
Research investigates human and AI attribute impacts on partially aligned human-AI interactions using 2,000 simulations and 290 human participants.
Research identifies consistent content selection biases in OpenAI, Anthropic, and Google LLMs, leading to polarization in content curation.
Research proposes Sequential Monte Carlo Speculative Decoding (SMCSD) to improve LLM inference speed by reweighting, rather than rejecting, draft tokens.
Research evaluates LLM-based agentic financial simulators (PersonaLedger) for generating differentially private synthetic data, finding fidelity in reproducing statistical distributions.
Research proposes a novel conformal prediction framework for LLMs using internal representations to improve uncertainty quantification beyond surface statistics.
Research claims stochastic tokenisation improves LLM robustness, reducing brittleness to adversarial attacks and input perturbations.
Research finds post-training reduces output diversity in language models, impacting inference methods and creative tasks.
Research identifies 'listener-speaker asymmetries' in LLM pragmatic competence, where models evaluate language differently than they generate it.
Research proposes RAGognizer, a method integrating a detection head during fine-tuning to reduce closed-domain hallucinations in RAG-augmented LLMs.
A new survey categorizes design principles and architectures for achieving intrinsic interpretability in large language models, contrasting with post-hoc methods.
Research explored token pruning to optimize multilingual LLMs (Qwen3, Gemma-3, Llama-3, Aya) for Korean-centric NLP, reducing size and improving efficiency.
Research finds LLMs (Gemini-Pro, GPT-4o Mini, Claude 3.7 Sonnet, DeepSeek-Chat, Llama 3) respond inconsistently to politeness across languages.
Research evaluates off-the-shelf LLMs as human surrogates in survey experiments, comparing their responses to human data for inferential consistency.
Research identifies hallucination in autoregressive models as early trajectory commitment due to asymmetric attractor dynamics, using same-prompt bifurcation on Qwen2.5-1.5B.
Researchers propose FineSteer, a unified framework for fine-grained inference-time steering in LLMs to reduce undesirable behaviors.
Research paper explores using anonymization techniques within Retrieval-Augmented Generation (RAG) pipelines to mitigate privacy risks in LLM applications.
Research proposes a faithfulness-aware uncertainty quantification method for RAG outputs to mitigate hallucinations arising from internal knowledge or retrieved context.
Research explores using multimodal LLMs to automatically detect misleading data visualizations by identifying violations of chart design principles.
Research identifies 'Miracle Steps' in LLM mathematical reasoning, where models achieve correct answers via unsound logic, showing reward hacking.
Research formalizes the 'one-sided conversation problem' (1SC), inferring missing speaker turns and generating summaries from single-party transcripts.
Research introduces MTR-DuplexBench, a new benchmark for evaluating full-duplex speech language models in multi-round conversations, addressing current single-round limitations.
Research proposes Next Token Probability Attribution (TPA) for detecting RAG hallucinations, accounting for all LLM components beyond context.
Research examines how LLMs resolve factual conflicts when retrieved information from different sources conflicts, focusing on source preference.
Research identifies 'new-knowledge-induced factual hallucinations' in LLMs after fine-tuning on new data, affecting previously known facts.
Research indicates LLMs assigned specific personas exhibit human-like motivated reasoning biases, mirroring identity protection in decision-making.
Research identifies prompt-induced hallucination mechanisms in Vision-Language Models (VLMs) for object counting, showing overstatement bias.
Comparative study evaluates Integrated Gradients, Attention Rollout, and SHAP for explainability on fine-tuned DistilBERT for sentiment analysis.
Research reviews training-free methods for enhancing LLM trustworthiness, covering hallucination, bias, toxicity, and adversarial robustness.
Researchers propose MemEvoBench, a benchmark to measure 'memory misevolution' in LLM agents, where contaminated memory leads to abnormal behavior.
Skill-RAG is a research paper proposing a RAG enhancement that uses LLM hidden-state probing to diagnose retrieval failure and dynamically route queries.
Research claims a data-efficient framework teaches reasoning models to code-switch, improving multilingual task performance without extra data.
New research proposes Sequential Internal Variance Representation (SIVR) to estimate LLM uncertainty from internal states to detect hallucinations.
Researchers introduced a new benchmark, the Metacognitive Monitoring Battery, to evaluate LLM self-monitoring across six cognitive domains using human psychometric methods.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion