Enhancing Rubric-based RL via Self-Distillation
Research explores self-distillation to enhance rubric-based Reinforcement Learning (RL) for LLMs, addressing limitations like limited exploration.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research explores self-distillation to enhance rubric-based Reinforcement Learning (RL) for LLMs, addressing limitations like limited exploration.
Research explores optimizing transformer architectures for specific datasets to understand optimal task-specific inductive biases beyond current scaling methods.
Researchers introduced ATLAS, a framework equipping Multimodal Large Language Models (MLLMs) with 'Think, Plan, Paint' reasoning for controlled image generation.
Research proposes Segment-Wise CoT Compression with Answer Alignment (SCA) to reduce LLM inference costs without compromising answer quality.
Research explores Differentiable Logic Gate Networks (Diff-Logic) for low-latency EEG classification on edge devices, compiling models into Boolean circuits.
Research proposes a Novel Hybrid Quantum Reservoir Computing (nHQRC) framework to detect phase transitions in non-equilibrium dynamical systems.
Research disentangles model (epistemic) and data (aleatoric) uncertainty in facial age estimation using Bayesian Neural Networks.
Research introduces RGMR, an inference-time framework adapting pre-trained foundation models for multi-scale temporal analysis and iterative refinement.
Research explores compact convolutional neural networks for detecting drones via their radio-frequency video transmissions, focusing on lightweight models.
ZifaMem proposes a structured memory system for AI companions, organizing dialogue into session summaries, episodic memories, and a consolidated user model.
Researchers propose Feature-Informed Diffusion Evolution (FIDE), a black-box framework for inverse rendering, bypassing differentiable renderers and gradient descent.
Research identifies a root cause for performance stagnation in PPO reinforcement learning, demonstrating that scaling to 1M parallel environments can prevent it.
Research introduces Time-Aware Prior Fitted Networks for zero-shot time series forecasting, leveraging exogenous variables to improve accuracy.
Research explores training-free interpretability methods for neural networks, questioning if expensive training-based methods yield superior insights.
Research identifies fundamental failures in marginal influence-based attribution methods for explaining global time series models due to computational mismatches.
Research indicates constant-stepsize SGD's behavior fundamentally changes for convex objectives with flat minima, deviating from Gaussian limits.
Researchers propose NeoST, a spatio-temporal foundation model trained entirely on synthetic data to avoid real-world data biases.
Researchers propose Overlapping Schwarz Attention, a hierarchical attention mechanism for LLMs inspired by domain decomposition methods, tested on operator learning.
Research introduces Jacobian-Aggregated Group Gradient (AGG) to reduce computational cost for Group Relative Policy Optimization (GRPO) in diffusion models.
FAIR-Calib proposes a Post-Training Quantization (PTQ) method to mitigate errors in Diffusion Large Language Models (dLLMs) caused by early, fragile token decisions.
HantaWatch proposes a federated learning framework for hantavirus genomic surveillance, allowing collaborative model training without raw data sharing.
Research evaluates machine learning models for Type 2 diabetes risk prediction, emphasizing external validation and fairness analysis on national populations.
ARGO, a smart eyewear platform, integrates on-device ML with an NPU (STM32N6 microcontroller) for low-latency, privacy-preserving local data processing.
A research paper proposes a predict-then-correct framework using few-shot continuous contextual bandits for adaptive retail demand forecasting.
A research paper proposes a Representation-Aware Distributionally Robust Optimization framework to improve model robustness against data shifts.
Research addresses the "factorization barrier" in diffusion language models, aiming to improve parallel token generation efficiency and coherence.
Research proposes 'Epiplexity per Joule' and 'Empowerment per Joule' metrics to connect AI's physical efficiency with intelligence.
Researchers propose TRACE, a method for safety patch learning to realign LLMs after fine-tuning without losing task utility.
Researchers developed a Bayesian framework, "diffusion-within-Gibbs sampling," for improved signal component decomposition in noisy data.
SOS-LoRA is a new parameter-efficient fine-tuning (PEFT) method improving upon LoRA by reparameterizing adapted weights to reduce interference.
A research survey identifies LLM unlearning as critical for cybersecurity, privacy, and safety by mitigating risks from memorized sensitive data.
Research explores hardware-aware design and optimization for edge intelligence systems, addressing complexity of deep learning on heterogeneous edge devices.
New research proposes Counterfactual Shapley Credit Assignment, a principled method to isolate policy skill from environmental stochasticity in RL agents.
Research explores how FFN residual writes influence retrieval accuracy in long-context models, showing native FFN scaling impacts state.
Research proposes Normalized Rewards to prevent over-optimization in Direct Alignment Algorithms (DAAs) like DPO, which can degrade LLM performance.
BACON proposes a four-stage pipeline for budgeted human calibration of AI judges to mitigate bias in model evaluation, ranking, and quality reporting.
Research explores OrthoGrad, a geometric intervention on optimizer updates, to reduce neural network memorization of noisy labels in training.
Research proposes a novel RL-guided genetic algorithm for multi-objective portfolio optimization, minimizing risk and maximizing return.
CyberGym-E2E is a new large-scale, real-world benchmark for evaluating AI agents' end-to-end cybersecurity vulnerability discovery and remediation.
Research explores mixed-timescale differential coding for federated learning model broadcast to reduce communication load in wireless environments.
Research introduces ECO, an efficient learning framework for Neural Combinatorial Optimization using batched preference optimization and a Mamba backbone.
DORA is an asynchronous reinforcement learning system that speeds up LLM post-training by overlapping generation with model training, addressing the rollout phase bottleneck.
A research paper proposes an 'AI epidemiology' framework to standardize expert-AI interaction data for prospective risk detection in deployed AI systems.
Research indicates LLM-generated CUDA kernels frequently employ 'reward hacking' to inflate performance against PyTorch on benchmarks like KernelBench, requiring co-evolving evaluation frameworks.
Research explores 'recombinatory-replay' in artificial memory, inspired by neuroscience, to drive insight and creative discovery across domains.
Research systematically investigates Reinforcement Learning (RL) based jailbreaking techniques in LLMs, highlighting threats to safe model deployment.
Researchers propose a multiverse-consensus pipeline for reproducible feature selection in untargeted LC-MS metabolomics, addressing pipeline decision variability.
Research proposes a low-bit KV-cache quantization method to reduce LLM inference memory and bandwidth costs while restoring accuracy.
A research paper proposes a volatility-aware machine learning approach to detect extreme price movements in high-frequency financial markets.
Research introduces a method for decomposing uncertainty in Bayes-filtered transformers, separating aleatoric from epistemic uncertainty for better risk assessment.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion