Anti-Periodic Positional Encoding: M\"obius Boundary Conditions Make In-Context Retrieval Reliable
New research proposes M"obius RoPE, an anti-periodic positional encoding method to improve in-context retrieval reliability in large language models.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
New research proposes M"obius RoPE, an anti-periodic positional encoding method to improve in-context retrieval reliability in large language models.
Researchers propose a multi-axis evaluation framework for structured audio captions, addressing limitations of existing metrics for multimodal outputs.
Research introduces TriviaRoomQA, a multilingual benchmark evaluating LLM performance on everyday, culturally grounded, and long-tail knowledge through quiz-style questions.
RUMBA introduces a new Russian-language benchmark for long-term conversational memory in LLMs, focusing on granular retrieval and reasoning.
Research identifies epanorthosis, a rhetorical self-correction, as a systematic overuse in LLM text due to training data and RLHF.
MedGame introduces an LLM-powered framework to transform static clinical cases into interactive, decision-centered storytelling games for medical education.
Research benchmarked five LLMs on multi-sensor physical hazard assessment, finding all consistently failed to issue precautionary warnings.
Research explores factors impacting the detection of deceptive outputs from LLMs, noting current probes fail in out-of-domain scenarios.
PersonaTrail introduces a new benchmark for personalized web agents, evaluating their ability to infer context from user browsing histories.
Research identifies "directional hallucinations" and ideological drift in LLMs when answering political questions, using a new measurement framework.
GenDB, a generative query engine using LLM agents, is demonstrated to automatically generate customized query processing code, aiming to reduce engineering effort.
Research introduces WaveformQA, a new benchmark to evaluate LLMs' temporal reasoning over digital waveform data, addressing a design verification gap.
NVIDIA-labs introduces Object-Oriented Agents (NOOA), a Python framework for building reliable AI agents by representing agents as Python objects.
Researchers propose "Refusal-Gated Decoding" to maintain LLM refusal behaviors when using high-temperature sampling for output diversity.
Research explores using transformer-assisted LLMs for source code summarisation to improve secure software development lifecycle maintenance.
Research introduces HiMe, a real-time, self-hosted, open-source personal agent platform for health insights from wearable data using LLM agents.
VibeVoice-ASR-BitNet introduces a highly compressed ASR model using INT8 and BitNet-style ternary quantization for edge CPU real-time inference.
Research explores open-weight LLMs for agentic coding on local, sensitive data, bypassing cloud transmission restrictions.
Research explores creating dynamic and physically realistic 4D virtual worlds from natural language using generative models, moving beyond manual graphics.
Research finds open-weight LLMs exhibit demographic disparities; Black-associated names lead to higher first-token entropy and more diverse continuations.
A new distillation framework, SCoRe, improves smaller LLM agents' multi-step reasoning by generating student-centric trajectories to narrow the performance gap.
Agentic Memory (AgeMem) proposes a unified framework for LLM agents to manage long-term and short-term memory, addressing context window limitations.
Research introduces DatedGPT, 1.3B-parameter LLMs pretrained on time-partitioned data to prevent lookahead bias in forecasting tasks.
LinearARD is a self-distillation method to restore original model capabilities when extending context windows with RoPE scaling, preventing performance degradation.
New benchmark, ImplicitBBQ, evaluates implicit bias in LLMs, addressing limitations of existing name-based proxies for detecting non-explicit identity biases.
FlyRoute, an arXiv paper, proposes a self-evolving framework for agent profiling that uses real traffic to dynamically update agent capabilities for task routing.
PennySynth uses RAG to generate quantum code for specialized frameworks like PennyLane, addressing LLM hallucination of specific quantum syntax.
Research introduces EvoSpec, an evolved speculative decoding method for LLM inference that adapts vocabulary and parameters in real-time to improve efficiency and quality.
MetaHOPE is a proposed framework for evaluating metaphor translation errors in machine translation and LLM outputs, focusing on semantic and cultural complexities.
WildTrace is a new benchmark for evaluating LLMs' ability to integrate evidence from disparate parts of long documents for complex reasoning tasks.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion