Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
Research reveals LLM resource allocation decisions shift unexpectedly based on context accumulated across deployment inference setups.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research reveals LLM resource allocation decisions shift unexpectedly based on context accumulated across deployment inference setups.
Researchers introduced SMTrap, an offline DoS attack method using SMT solvers to trigger heavy search compute in large reasoning models.
Research introduces Dual-Bounded Relational Recall, combining top-k seeds with graph-adjacent context to boost RAG without added budget.
Researchers introduced BudgetDoc and DRB, a 1B multimodal router that predicts LLM compute requirements for complex document tasks.
Research proposes self-evolving metric frameworks using small Python defect-detectors to evaluate complex outputs like generated reports.
ArXiv paper proposes a framework for enterprise AI agents to selectively unlearn obsolete policies, regulations, and workflow rules.
Research demonstrates that standard Gaussian noise defenses fail to prevent text inversion attacks on dense vector embeddings.
CDSP pipeline converts investment committee transcripts into structured predictive signals via LLM context labeling and taxonomy mapping.
New research shows aggregate benchmark scores obscure specialized model strengths on tabular datasets, proposing peak performance metrics.
Paper demonstrates that updating agent harnesses—prompts, tools, and routing—causes behavioral drift even when underlying LLM weights are frozen.
An analysis of 500 Hugging Face model cards reveals current documentation standards fail to support downstream governance of open-weight models.
Researchers introduced FraudBench, a benchmark designed to stress-test policy-grounded conversational banking agents against adaptive fraud attacks.
OpenAI published a model card for its compact, bidirectional Privacy Filter designed for single-pass PII and secrets redaction.
Research analyzes static task properties that drive difficulty in software issue resolution benchmarks for AI coding agents.
Researchers introduced SESSE, a training-free framework decomposing LLM-as-judge evaluations into structured sub-questions.
Research demonstrates post-training 4B models to self-restrict task authority in terminal and Model Context Protocol agent environments.
New research shows bitsandbytes post-training quantization (INT8/INT4) worsens memory retrieval degradation in LLMs over long conversations.
An arXiv paper argues AI evaluation must pivot from capability to precision—measuring output variance across identical requests.
Paper reveals supervision targets in deep financial forecasting models should systematically deviate from actual inference horizons.
Researchers introduced NINJA, a jailbreak method hiding harmful goals in long-context inputs using benign model-generated content.
Diagnostic evaluation of AI agents on 100 frontier research tasks identifies systematic operational failure modes in autonomous workflows.
Researchers propose a constrained exploration-exploitation process to prevent LLM agents from overfitting during self-evolution and skill optimization.
WorldPack introduces a dynamic frame compression method for long-context video world models, addressing consistency in future visual generation.
Researchers demonstrated full-parameter post-training of trillion-parameter-scale MoE models on Ascend NPU SuperPOD, optimizing for memory, communication, and kernel efficiency.
New research argues that current methods for evaluating LLM activation explanations are structurally flawed, failing to penalize individual false claims.
Research identifies 'phantom transitions' in LLM fine-tuning where cross-entropy loss decreases but correct token ranking fails to improve.
Research demonstrates LLM-driven AutoML using GPT-5, GPT-4o, and Claude Sonnet 4 to autonomously design and refine neural architectures for cross-lingual handwritten OCR.
Research introduces "Diversion Decoding," a method to improve hallucination detection in LLMs by identifying factually incorrect generations.
Research presents methods to predict RAG performance gain for question answering, identifying a novel post-generation predictor as most effective.
Stripe acquired OpenRouter, an AI prompt-routing platform, to strengthen its infrastructure positioning for agentic workflows and AI payments.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion