Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
Research proposes a logit biasing method for adapting black-box language models for domain-specific tasks and privacy without full fine-tuning.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research proposes a logit biasing method for adapting black-box language models for domain-specific tasks and privacy without full fine-tuning.
New research proposes Direct Corpus Interaction (DCI) for agentic search, guiding fine-grained corpus exploration with relevance estimates.
Researchers propose SINT-Flow, an LLM-based framework with five operators for fully automated, end-to-end schema integration across diverse input tables.
Research introduces XIH-Bench, a new benchmark to evaluate instruction hierarchy compliance in multilingual LLMs, finding language impacts adherence.
Research introduces HCG-RAG, a method using schema-constrained causal graphs for retrieval-augmented generation to reduce graph size and cost.
Research compares LLM-based vs. lexicon-based sentiment analysis for detecting tail-risk signals from Reddit data on meme stocks like GME and AMC.
Researchers propose Zhijing, a framework for measuring and integrating social intelligence in LLMs, including the SoMBench benchmark.
Research identifies two distinct failure points in compressed short-text generation: information loss in the codec or weak codes from the latent generator.
Research finds LLM prompt tone significantly impacts inference cost via output token length, with less effect on accuracy across seven tones on MMLU.
Researchers propose Mixture of Language Group Experts (MoLGE) to improve performance and efficiency in massively multilingual automatic speech recognition models, addressing 'curse of multilinguality'.
SyRuP, a new method, aims to improve LLM adherence to complex system prompts during decoding without requiring model tuning or reranking.
Research proposes a 'frozen model' architecture with a persistent memory of verified solutions for 100% accuracy and zero inference tokens.
Research quantifies the 'tokenizer tax' for Indian languages, showing LLMs incur higher processing costs due to English-centric subword tokenizers.
Earnings25, a new 500-hour finance-domain speech benchmark for ASR evaluation on English earnings calls, is introduced by arXiv research.
Research finds appending a two-word confirmation tag to a question changes LLM responses, measuring this 'tag effect' across 45 models.
Research finds adapter merging in two-stage fine-tuned LLMs can reactivate latent reasoning traces, even causing interference post-alignment.
CONSISTRE is a new framework designed to improve consistency in document-level relation extraction using large language models by enforcing relational constraints.
New research introduces a two-layer evaluation framework to separate language model execution evidence from correctness, exposing how models fail beyond accuracy scores.
New research from arXiv proposes Cross-Attention Calibrated Deduplication to improve RAG system efficiency by identifying and removing redundant data chunks.
Researchers developed formally verified IEEE-754 FP32 and BF16 arithmetic for ARCH HDL, a language intended for LLM generation, ensuring mathematical correctness.
Research indicates self-improving agents using self-authored tests for verification can show high internal scores while real performance degrades.
Research identifies a 'blind spot' in AI agent long-term memory systems where retrievers fail to link implicit knowledge to queries.
INS-ActBench, a new benchmark, evaluates LLMs on professional actuarial tasks requiring auditable, context-grounded, and tool-executable decisions.
Research evaluates closed-loop validation-repair for clinical LLMs in healthcare to achieve structured output schema compliance (ICD-10, CPT, HL7 FHIR).
Research introduces Cognitive Attribution Graphs (CAGE) to improve inline citation generation in long-form LLM outputs, addressing 'attribution ambiguity'.
RM-Distiller proposes using generative LLMs more effectively for reward model distillation, moving beyond simple binary annotation.
Research explores sparse autoencoders (SAEs) to better link SAE features to LLM behavior, addressing inconsistencies in causal effects and steering.
Research finds LLMs can generate highly personalized disinformation across languages and demographics, challenging existing safety safeguards.
Research identifies a gap in conjunctive cross-page retrieval, where current systems struggle to confirm all evidence for multi-part requests within a document.
Research reveals score-conditioned In-Context Learning (ICL) in LLMs corresponds structurally to policy gradient optimization, explaining iterative output improvements.
Research identifies a fundamental limitation in dual-encoder vision-language models where compositional queries like "A and not B" fail due to a Bag-of-Concepts effect.
Research evaluates how model scale and quantization affect uncertainty signals in Vision-Language Models (VLMs) under image degradation, impacting confidence for deferral decisions.
Research finds behavioral detection of unfaithful Chain-of-Thought (CoT) explanations fails when LLM answers are incorrect, hindering oversight.
Research finds fine-tuning small models with legal context improves their accuracy on legal Q&A, even with retrieved law.
Research explores how task-adaptation methods like supervised fine-tuning (SFT) impact large language model alignment, including safety and other behaviors.
New arXiv research proposes a controlled environment and distillation method to improve multi-turn long-horizon planning in foundation model agents.
LA-RL introduces a label-aware self-reflection method for reinforcement learning in information extraction, improving structured output correction.
Research explores Parallel Autoregressive Decoding (PARD) for block diffusion language models, showing better alignment with left-to-right generation.
Research shows frontier LLMs use 'invisible reasoning' with semantically irrelevant filler tokens to improve performance on synthetic tasks.
Researchers introduced LEX-EC, a black-box audit framework for zero-shot LLM personality classification, using lexical ablation to distinguish signal from marginal-distribution effects.
Research explores pointer-augmented autoregressive generation for patent claims, addressing hierarchical constraints in structured text with LLMs.
Research indicates LLM debate patterns differ across languages, with Chinese models less prone to repeating arguments compared to other languages tested.
Research on Multi-head Latent Attention (MLA) in DeepSeek-V2 details how it compresses key-value pairs, reducing KV-cache by 81% during inference.
Researchers introduce Tokengeist, a multi-turn attribution tracing method for agentic conversations, improving lineage tracking for LLM responses.
ELMOD, a 2.7B parameter German language model, is introduced for efficient on-device inference using publicly available data and optimized preprocessing.
Kalypso introduces relational LLM serving, an abstraction making LLM execution aware of query plans to improve performance for semantic operations on unstructured data.
Research highlights unaddressed security and governance risks of third-party API routers in agentic AI workflows, especially regarding data inspection and modification.
Researchers find frontier LLMs exhibit systematic failure in predicting the outcomes of computational processes and algorithm executions.
Research proposes Source-Aware Reranking for RAG, incorporating source provenance and credibility as a reliability prior in document retrieval.
Research proposes a method to audit LLM alignment controllability, measuring how far responses can be steered from a 'resting point' via system prompts.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion