FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
Research scales synthetic instruction-tuning data to pre-training levels, aiming to overcome limited supervised data for LLMs.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research scales synthetic instruction-tuning data to pre-training levels, aiming to overcome limited supervised data for LLMs.
DR-Arena introduces an automated evaluation framework for Deep Research (DR) Agents, addressing limitations of static benchmarks for LLM-based autonomous investigation.
Research introduces Engram, a conditional memory module for LLMs, aiming to improve efficiency by adding a new sparsity axis for knowledge lookup.
BiasLab introduces a multilingual dual-framing framework for measuring LLM bias, with specific application to workplace and HR contexts.
Research introduces destroR, a benchmark and adversarial-training defense pipeline for Bangla transformer models against meaning-preserving attacks.
PiCSAR is a training-free method improving LLM reasoning by probabilistically selecting and ranking best-of-n candidate solutions without ground truth.
REAL is a new framework for precisely identifying behavior-relevant internal modules (attention heads/layers) in LLMs for inference-time steering.
Research identifies a two-stage training dynamic in Transformers, where models like GPT-2 progress from syntactic to semantic correctness.
Research details a mechanistic interpretability approach to identifying and mitigating scoring biases in LLM-as-judge applications at the representation level.
Research explores theoretical understanding of Transformer models, focusing on expressivity and sample complexity via C-RASP for narrower teachers.
Research investigates if 'lesioning' multimodal language models can reproduce human aphasia-like picture-naming errors, simulating brain injury.
SCOPE-RL proposes a reinforcement learning framework to optimize LLM reasoning paths by providing feedback before and after verifiable success.
Research finds Vision-Language Models (Qwen3-VL-30B-A3B, LLaVA-1.5-13B) exhibit significant context-dependent affordance drift.
Researchers introduced LMEB, a benchmark for evaluating memory embeddings in long-horizon retrieval tasks, addressing gaps in current text embedding evaluations.
A new benchmark, IslamicMMLU, was introduced to evaluate LLMs on Islamic knowledge with 10,013 multiple-choice questions across Quran, Hadith, and Fiqh.
Research introduces StructAgent, a digital agent framework designed for long-horizon tasks using unified causal structures to improve task progress interpretation.
Research introduces StanceMoE, a Mixture-of-Experts (MoE) architecture to improve actor-level stance detection by capturing heterogeneous linguistic signals.
Research introduces a methodology to enhance Retrieval Augmented Generation (RAG) systems by integrating an auxiliary feedback RAG for human feedback.
Research investigates how human annotator competence evolves over time in subjective tasks like social influence recognition, involving experts across groups.
Researchers propose Tool-MCoT, a small language model (SLM) fine-tuned with external tools for multimodal content safety moderation, addressing LLM cost and latency.
Research introduces GCSR, a generative framework for Chinese statute retrieval, reformulating the task as sequence generation to bridge query-language gaps.
Research evaluates LLMLingua-2, a prompt compression technique, on diffusion large language models (DLLMs) like LLaDA-8B-Instruct across reasoning and summarization tasks.
Research on 'domain-aware scaling laws' explores how combining data from different domains influences large language model performance, identifying synergies and interferences.
Researchers propose Next Implicit Token Prediction (NITP), a pre-training method augmenting discrete next-token prediction to improve LLM latent representations.
Research finds AI-generated ideas from five agent frameworks and LLMs narrow scientific exploration, reinforcing existing work rather than broadening it.
PerspectiveGap, a new benchmark, evaluates LLMs' ability to create orchestration prompts for multi-agent systems, focusing on sub-agent knowledge.
Research paper introduces LOGOS, a pluggable layer for self-evolution and governance in multi-agent AI systems, to control agent behavior.
Research benchmarks LLMs against urban planners for contextual sensitivity, value awareness, and institutional literacy in professional judgment.
Research introduces KVEraser, a method to efficiently remove specific context from LLM KV caches post-hoc, addressing stale or incorrect information.
Research introduces STEC, a method for evidence compression in LLM-based multi-hop question answering to improve final answer selection.
Research proposes SelfCompact, a method for LLM agents to autonomously manage and compact their context window based on task structure, reducing stale content.
Research finds that standard post-training methods (SFT, RL) can degrade pre-trained 'compassion' values in a Llama 3.1 8B model depending on post-training data domain.
Research explores using LLM reasoning traces to predict human item difficulty in educational assessments, focusing on cognitive processes.
A survey proposes shifting AI-generated video detection from artifact-centric methods to high-level semantic verification and factual fidelity.
Research applies multiple instance learning to automate cancer registry tumor classification using patient-level labels, addressing data scarcity.
Research presents SAMPA, a Whisper-based model for automatic prosodic boundary segmentation in Brazilian Portuguese speech, improving on rule-based methods.
XALPHA proposes a memory-driven AI quant researcher for end-to-end alpha discovery from hypothesis to code, leveraging LLMs.
Research views self-attention in Transformers as a connection walk, identifying it as a specific operator on token-position graphs.
Soofi S 30B-A3B, a sovereign, open-source MoE hybrid Mamba Transformer model for German and English, claims high throughput for long context.
SpurLens detects spurious correlations in Multimodal LLMs using GPT-4 and object detectors, revealing inherent biases.
Research paper argues LLM-based social simulations need boundaries due to their tendency to produce homogeneous, 'average persona' outputs, limiting diversity.
Research introduces a constraint-aware hierarchical search method for fine-grained classification tasks driven by regulatory rules, not just semantic similarity.
Research compares Vision-Language Models (VLMs) across domains, highlighting that benchmark performance does not predict real-world behavior across varied datasets.
New benchmark SWE-MERA aims to dynamically evaluate LLMs on software engineering tasks, addressing data contamination issues in SWE-bench.
CRINN introduces a contrastive reinforcement learning method to optimize Approximate Nearest-Neighbor Search (ANNS) algorithms, targeting execution speed.
Nested-ReFT proposes an efficient reinforcement learning method for large language model fine-tuning (ReFT) using off-policy rollouts to improve reasoning.
Research addresses emotion recognition in signers, creating new datasets for Japanese Sign Language and British Sign Language to overcome data scarcity and grammatical-affective expression overlap.
TagSpeech presents a unified LLM-based framework for end-to-end multi-speaker ASR and diarization with fine-grained temporal grounding.
MUGEN benchmark reveals large audio-language models (LALMs) struggle with multi-audio understanding, especially with increased concurrent inputs.
Research introduces RecursiveMAS, a framework extending recursive computation from single LLMs to multi-agent systems for deepened reasoning.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion