Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D
Frontier LLMs struggle with exact string copying due to Transformer positional encodings, indicating limitations in fundamental operations.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Frontier LLMs struggle with exact string copying due to Transformer positional encodings, indicating limitations in fundamental operations.
Research identifies 'shortcut learning' in transformer-based speech and language models, where models rely on spurious correlations.
Research compares token, byte, and pixel encodings for language models across 13 languages, controlling linguistic content and model capacity.
ToolSciVer is a new research model that uses visual tool-augmented reinforcement learning for multimodal scientific claim verification.
Research redefines empathy in AI as predictive misalignment tolerance and proposes a co-regulation framework for dialogue repair.
Research demonstrates that malicious chain-of-thought (CoT) traces can be transferred between LLMs, inducing harmful behavior in target models.
Research proposes Looped Latent Attention (LLA) to compress K/V caches in looped Transformers, reducing memory footprint and potentially inference cost.
CoWeaver is an arXiv research paper proposing a bidirectional, learnable, and explainable algorithm for matching human scientists and AI agents for collaboration.
Research finds AI watermarks, increasingly mandated by regulations like the EU AI Act, are not forensically reliable for legal evidence.
Research proposes HCIG, a Hierarchical Cross-Modal Incongruity Graph Network for detecting multimodal sarcasm and cyberbullying.
Research introduces ActiveVision, a new benchmark to test if multimodal LLMs use active observation in vision tasks, mimicking human visual processing.
Research introduces Mixture-of-Neighbors Induction Memory, enabling semiparametric language models to learn and update non-parametric memory dynamically.
Researchers demonstrated Latent Fusion Jailbreak (LFJ), a white-box attack manipulating LLM internal representations to bypass safety alignments.
Research explored moral biases (Knobe effect) in finetuned LLMs using mechanistic interpretability and Layer-Patching across three open-weight models.
New research proposes Bifocal Attention, an improvement over Rotary Positional Embeddings (RoPE) to better capture long-range, periodic structures in LLMs.
Research details a method for generating high-quality, long-horizon terminal interaction data using Docker for training agentic models.
Research explores 'speculative decoding with a speculative vocabulary' to accelerate large language model inference using a smaller draft model.
Research introduces RLearner-LLM with Hybrid-DPO to address the logical alignment gap in DPO, improving factual correctness over fluency bias.
PrimeFacts is a new methodology and resource for extracting fine-grained evidence from fact-checking articles for automated verification systems.
Researchers propose VCG-Bench, a new benchmark and "Diagram-as-Code" paradigm for Vision-Language Models to improve structured diagram generation and editing.
Research indicates LLMs encode syntactic distinctions beyond Universal Dependencies, specifically around finite and infinitival clauses in wh-movement.
Research explores the cost-quality trade-offs in 'skill rewriting' for LLM agents, finding shorter skills can increase agent operational costs.
Research demonstrates that large language model (LLM) alignment guardrails can be broken with one-shot Group Relative Policy Optimization (GRPO) using a single biased example.
AuAu is a new research benchmark designed to assess and quantify authoritarian tendencies and alignment in large language models using psychometric and vignette-based tests.
Research analyzes how deep transformers, dominant in language modeling, form hierarchical representations across layers for expressive power.
Research finds LLM-as-a-judge scores are unreliable for optimizing closed-loop table recognition, showing weak signals and non-reproducible rankings.
Research uses linear probing and Bloom's Taxonomy to mechanistically interpret cognitive complexity in LLMs via internal neural representations.
Researchers introduced Jailbreak Foundry (JBF), a system translating LLM jailbreak papers into executable modules for unified evaluation.
Research paper identifies 'AI-Fiction Paradox': LLMs trained on fiction data struggle to generate compelling long-form fiction.
Research proposes Contrastive Hypothesis Retrieval to improve RAG for medical Q&A by explicitly ruling out hard negatives semantically close but clinically distinct.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion