Twin Agent: Context Residual Compression for Privilege Separated Agents
Research introduces 'Twin Agent', a method for secure LLM agents using context residual compression to mitigate prompt injection risks.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces 'Twin Agent', a method for secure LLM agents using context residual compression to mitigate prompt injection risks.
Researchers propose PoTRE, a framework using multiple AI agents for complex reasoning to address LLM struggles with long-horizon planning and error correction.
Spectral-LSH, a new training-free method, offers sub-quadratic prompt compression by approximating dominant attention components before LLM inference.
Research addresses 'scaffolding collapse' in LLM-based Socratic tutors, where models abandon guided inquiry for direct solutions under pressure.
Researchers propose Janus, a foresight framework using multi-agent simulation to anticipate delayed operational risks from tool-using agents.
Research introduces Learn2Discern (L2D), a new framework and benchmark to evaluate LLM information discernment from external knowledge sources based on normative axioms.
Research explores if Large Language Models (LLMs) can benefit from “experiential abstractions,” similar to humans distilling experience into reusable strategies.
Research introduces AdaRoPE, optimizing Transformer positional embeddings by adapting frequency schedules and scaling for different attention heads.
Research finds autointerpretability scores for sparse autoencoders are significantly influenced by evaluation pipeline choices, not just feature properties.
Research evaluates 21 LLMs' ability to recognize human values based on Schwartz's ten basic values using 1,000 Russian situational texts.
Research systematically investigates challenges of culturally loaded machine translation using 'Dream of the Red Chamber' to evaluate LLM performance.
PyroDash introduces a cost-aware framework enabling token-level collaboration between small and large language models, reducing inference costs.
Researchers introduced the Maskability Index (MI), a quantitative metric to predict optimal prompting strategies for language models.
Research identifies a benchmark leakage issue where 'out-of-distribution' classes were inadvertently included in training, skewing OOD detection metrics.
Research finds generative AI produces book-length fiction at near-zero cost, impacting the self-published market on Amazon from 2023-2026.
A new research benchmark, HalluTruthQA, is introduced for fine-grained hallucination detection, localization, and explanation in Arabic LLM outputs.
New research, HyGRL, proposes a hybrid graph reasoning framework for multi-entity questions, addressing limitations in conventional RAG and LLM-constructed Graph-RAG.
Research demonstrates LLM stance sensitivity to linguistic construction, not just lexical choice, using causal tracing to localize shifts within models.
Researchers demonstrated full-parameter post-training of trillion-parameter-scale MoE models on Ascend NPU SuperPOD, optimizing for memory, communication, and kernel efficiency.
Research identifies three distinct modes of sycophancy in large language models, challenging the assumption of it being a single dimension.
New research introduces OpenSkillRisk, a benchmark for evaluating the safety of LLM-based agents that use real-world, risky third-party skills.
Researchers propose a novel framework using Clopper-Pearson confidence intervals to compute rigorous probabilistic safety bounds for LLM harmful output.
Researchers propose a two-process psychometric theory for evaluating language model self-reports, addressing current ad hoc prompting and validation issues.
Research finds standard knowledge distillation (KD) harms over half of training samples for low-resource language summarization, yielding minimal ROUGE-L improvement.
Research explores using reinforcement learning to enable retrieval-augmented LLMs to selectively adopt valid evidence while rejecting misleading information from retrieval results.
New research proposes methods for LLM preference alignment that reward reasoning trajectories, not just final outcomes, addressing coarse credit assignment.
Researchers propose a statistically-grounded method for sparse-feature intervention in LLMs, enhancing activation steering for behavioral control without fine-tuning.
Researchers propose Hibiki-Zero, a new method for simultaneous speech-to-speech translation that eliminates the need for word-level aligned data.
Researchers introduced Sentence Splitter, a self-supervised T5-based model designed to uncover latent factual structures within natural language sentences.
Research introduces specialist language models for high-precision, scalable missing value prediction in tabular data, addressing overconfidence and hallucination.
DocOps introduces a deterministically verifiable evaluation framework for autonomous agents focused on complex document operations, deconstructing tasks into atomic dimensions.
Research demonstrates materials science mechanism information is readable and steerable in google/gemma-4-E4B-it's hidden states, enabling controlled transformations.
Researchers propose JailMeter, a new evidence-based framework to consistently evaluate jailbreak attack effectiveness on large language models.
New research proposes "knowledge-centric self-improvement" for AI agents, shifting focus from optimizing agent design to improving knowledge bases for better transferability and maintenance.
Research identifies a 'structural trilemma' in LLM responses within emotionally sensitive contexts, risking maladaptive reinforcement or over-restriction.
Research finds small language models struggle to follow instructions when they conflict with their inherent task behaviors across MCQA, sentiment, and math tasks.
FinMMEval 2026 Task 1 introduces a benchmark for multilingual financial multiple-choice QA across English, Chinese, Arabic, and Hindi.
Research explores using hypernetworks for train-time knowledge injection into LLMs by generating LoRA adapters for large factual corpora.
SLPO (Scaling Latent Reasoning via a Surrogate Policy) proposes a method to improve latent reasoning in LLMs, which uses continuous vectors for intermediate computation, potentially outperforming explicit Chain-of-Thought.
Researchers propose a reference-free framework for evaluating reasoning in LLM-generated open-ended answers by decomposing reasoning traces into segments and using Natural Language Inference (NLI).
Research finds Supervised Fine-Tuning (SFT) in LLMs reduces behavioral diversity in sequential decision-making tasks, potentially narrowing reasoning.
Research explores Chain-of-Modality reasoning for spoken language models to improve mathematical question answering over verbalized expressions.
New research proposes rubric-oriented document set selection and ranking for AI agents, moving beyond relevance to address inter-document interactions.
VizRAG introduces hypergraph visualization to Retrieval-Augmented Generation, enhancing multimodal LLMs beyond text-centric knowledge retrieval.
D2VBench introduces a new benchmark with 10,000 daily-scenario value dilemmas to improve evaluation of LLM value alignment, addressing limitations of existing benchmarks.
TriAgent, a multi-agent committee, improves cost-efficiency for financial sentiment analysis by stratifying queries across models of varying complexity.
Research identifies 'contextual entrainment' in unimodal language models and proposes ENTRAP-VL to investigate its manifestation in vision-language models.
FinMMEval 2026 Task 2 introduces a multilingual financial short-answer question answering benchmark for LLMs using financial statements and news.
Research proposes a Conversational Risk Accumulation (CRA) framework with stateful guardrails to address safety failures in multi-turn LLM systems.
Research finds verbatim conversational chunks outperform LLM-extracted structured artifacts for long-conversation memory recall in retrieval-rerank-reasoning pipelines.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion