Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
Research tests how logical judgments in LLMs respond to learned contextual soft prefixes, diagnosing syllogistic stability.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research tests how logical judgments in LLMs respond to learned contextual soft prefixes, diagnosing syllogistic stability.
Octopus introduces fine-tuning for 2B, 3B, and 7B parameter on-device LLMs to enhance software API function calling using a new dataset.
Octopus v3 introduces a sub-billion parameter multimodal AI agent designed for on-device inference, processing language, visual, and audio inputs.
Octopus v4 research introduces a 'Graph of Language Models' to orchestrate multiple niche LLMs, aiming for more efficient and capable systems.
Researchers propose Octo-planner, an efficient on-device Planner-Action framework that separates planning and action execution for AI agents.
Research introduces Dolphin, a decoder-decoder architecture using a 0.5B parameter model to distill long contexts for a 7B LLM, improving energy efficiency.
A new arXiv survey reviews Knowledge-Oriented Retrieval-Augmented Generation (RAG) techniques, enhancing LLM accuracy with external knowledge.
Research proposes Finetuning-aligned Sequential Training (FAST) for Sparse Autoencoders (SAEs) to address destructive gradient noise in instruct models.
New benchmark, PapersPlease, evaluates LLM decision-making in 3,700 moral dilemmas, focusing on prioritizing human needs based on ERG Theory.
New research proposes PACS, an Implicit Actor Critic Coupling framework, to improve Reinforcement Learning with Verifiable Rewards (RLVR) for LLMs on reasoning tasks.
Research proposes abductive reasoning for generating executable interaction trajectories to overcome data scarcity in RAG agent development.
Research indicates LLMs implicitly encode problem difficulty in their internal representations, detectable via linear probing.
Research explores fine-tuning LLMs for in-parameter graph reasoning, aiming to overcome token overhead of text-converted graphs.
Research addresses VLM hardware-model mismatch on NPUs, identifying quantization brittleness and I/O-bound autoregressive attention as key issues.
Research proposes Supervised Moral Rationale Attention (SMRA), a self-explaining hate speech detection framework using moral rationales for attention alignment.
Research identifies 'sockpuppetting' as an enhanced prefill attack, improving LLM jailbreaking through ensembled prefill variants.
Research reveals mechanistic insights into how backdoor attacks operate in LLMs, specifically identifying trigger formation and processing within attention heads.
Research details MiLMMT-46, a Gemma3-based open LLM adapted for multilingual machine translation across 46 languages through continual pretraining and finetuning.
Research demonstrates a method for systematically identifying rare, critical safety failures in large language models by efficiently sampling diverse responses.
Research identifies 'truncation blind spot' in LLM decoding (top-k, nucleus sampling), systematically excluding human-like, low-probability token choices.
Research proposes 'retromorphic testing with hierarchical verification' to detect hallucinations in RAG by judging faithfulness against retrieved context.
Research investigates why LLMs exhibit overconfidence in incorrect answers, identifying internal mechanisms that cause inflated verbalized confidence.
ClawBench introduces an evaluation framework with 153 real-world online tasks across 144 platforms to test AI agents' ability to complete web-based workflows.
Research argues current LLM multilingualism is incidental, leading to unequal, brittle, and opaque behavior across languages due to training on uneven corpora.
Research explores context compliance in RAG, specifically whether models follow retrieved evidence versus prior knowledge under conflicting information.
Research explores using LLMs to generate synthetic, taxonomy-targeted errors for quantitative reasoning tasks, aiming to aid tutoring and education.
Research proposes method to improve answer extraction and generation in context-based question answering systems using LLMs, addressing ambiguity and consistency.
Research finds language models with "brittle memory" can re-emit stale, incorrect information confidently if memory retains conclusions but loses supporting work.
New benchmark, MTEB-PT, introduced for evaluating Portuguese text embedding models across 14 datasets, addressing underrepresentation in current metrics.
Research introduces MAESTRO, a method for structured pruning of Mixture-of-Experts (MoE) models to reduce memory footprint and improve deployment efficiency.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion