HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
Research introduces HiMe, a real-time, self-hosted, open-source personal agent platform for health insights from wearable data using LLM agents.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research introduces HiMe, a real-time, self-hosted, open-source personal agent platform for health insights from wearable data using LLM agents.
New research, "Moir," proposes a knowledge editing method for LLMs that mitigates degradation of reasoning capabilities by letting the model direct its own edits.
Researchers propose "Refusal-Gated Decoding" to maintain LLM refusal behaviors when using high-temperature sampling for output diversity.
NVIDIA-labs introduces Object-Oriented Agents (NOOA), a Python framework for building reliable AI agents by representing agents as Python objects.
New benchmark, ImplicitBBQ, evaluates implicit bias in LLMs, addressing limitations of existing name-based proxies for detecting non-explicit identity biases.
Research explores methods to make watermarks in open-source LLMs durable against post-training modifications like model merging.
Research identifies 'Routing Subspaces' to locate why fine-tuned LLMs may appear safe in evaluation but still exhibit problematic behaviors in use.
MetaHOPE is a proposed framework for evaluating metaphor translation errors in machine translation and LLM outputs, focusing on semantic and cultural complexities.
LinearARD is a self-distillation method to restore original model capabilities when extending context windows with RoPE scaling, preventing performance degradation.
OpenForgeRL is an open-source framework designed to train AI agents that use complex, multi-turn inference harnesses like Claude Code or Codex.
Research introduces DatedGPT, 1.3B-parameter LLMs pretrained on time-partitioned data to prevent lookahead bias in forecasting tasks.
Research explores if Vision-Language Models (VLMs) express visual content with discourse-appropriate information structure, using Hungarian language testing.
New research proposes "Answer-then-Edit" for anti-distillation, generating defensive LLM outputs to prevent unauthorized knowledge extraction while preserving utility.
Research introduces TopoGuard, a graph theory-based defense against split-knowledge attacks on RAG systems where individually benign documents create false associations.
A new distillation framework, SCoRe, improves smaller LLM agents' multi-step reasoning by generating student-centric trajectories to narrow the performance gap.
Research explores creating dynamic and physically realistic 4D virtual worlds from natural language using generative models, moving beyond manual graphics.
RE-AD framework uses LLMs to proactively validate human data labeling quality in real-time, improving annotation accuracy for model training.
Research finds all frontier LLMs exhibit response drift, producing outputs that deviate from expert-validated references, uncharacterised by human evaluation.
Research paper introduces Semantic Field Theory (SFT) as a computational model for lexical semantics and stabilized interpretation, refining its mathematical core.
Research explores how inherent narrative structures and archetypal roles in LLM training data systematically influence model behavior and pose governance risks.
New research proposes THOR, a Theta-Gamma hierarchical oscillatory reasoning framework, to improve multi-hop question answering by addressing attention decay and error accumulation.
CAMeR introduces a keyword-gated hybrid activation memory framework for LLM agents to selectively retain relevant information over extended dialogues.
Instruct-FD introduces an instruction-conditioned benchmark for evaluating controllable turn-taking in full-duplex spoken dialogue systems.
Pulsar Attention proposes a new method for distributed LLM inference that replaces static context anchors with content-aware components, reducing compute costs.
Research tested GPT-4.1's ability to predict opinions using persona simulation, accurately forecasting 2024 election outcomes based on U.S. state personas.
New arXiv research introduces Capital Markets LLM Reliability Score (CM-LRS) for evaluating LLMs on "bankability" in capital markets workflows.
A preliminary research study explores whether valence (emotional tone) in natural language can reflect morality, impacting AI ethics.
Research fine-tunes small language models (0.6B-20B parameters) to translate natural language into MiniZinc, a domain-specific constraint language.
Research investigates the effectiveness of LLMs in detecting their own generated content across programming and writing tasks.
Researchers propose Constrained Shared-Private Fusion (CSPF), a new method for more reliably evaluating non-verifiable AI tasks by integrating diverse reward models.
REFACT proposes an adaptive method for LLMs to restate facts in their chain-of-thought reasoning, aiming for compact, faithful, and context-aligned outputs.
Researchers introduced Rushes, a human preference dataset for pluralistic alignment collected from interactive AI-generated narratives.
GenDB, a generative query engine using LLM agents, is demonstrated to automatically generate customized query processing code, aiming to reduce engineering effort.
Research introduces WaveformQA, a new benchmark to evaluate LLMs' temporal reasoning over digital waveform data, addressing a design verification gap.
Telco-GAIA introduces a bilingual, multi-modal benchmark for evaluating tool-using agents on real-world telecom data with complex reasoning.
Research finds open-weight LLMs exhibit demographic disparities; Black-associated names lead to higher first-token entropy and more diverse continuations.
Research introduces REGARD to study affective framing differences in LLMs, moving beyond simple sentiment to understand regional biases.
Research explores methods like low-rank decomposition and quantization to compress large language models, aiming to mitigate performance degradation at high ratios.
Research finds LLMs can generate deceptive responses with high confidence, increasing their persuasiveness to users, across various models and datasets.
TextGrad, a method for optimizing language model text components from natural language feedback, struggles with agentic systems due to delayed feedback attribution.
Research explores open-weight LLMs for agentic coding on local, sensitive data, bypassing cloud transmission restrictions.
Researchers introduced TOPReward, a training-free method using LLM token probabilities as zero-shot rewards for general-purpose robot learning, reducing manual effort.
Research demonstrates scaling closed-loop LLM-based channel configuration for neural networks by optimizing widths via executable code generation.
Research introduces 'swarm-attack,' an open-source adversarial framework using coordinating LLM agents to bypass safety and discover software vulnerabilities.
Research introduces CT-Merging, a method to combine multiple LoRA adapters into one multi-task adapter, reducing storage and inference complexity.
Research proposes sparse autoencoders (SAEs) to create interpretable embeddings for analyzing large text corpora, addressing limitations of LLM-based and dense embedding methods.
Chronofy, a new RAG architecture, introduces a temporal-logical decay mechanism to prevent LLMs from using obsolete facts, addressing 'temporal hallucination'.
Research identifies 'attention specific heat' as a reliable predictor for Grokking, improving neural network generalization and training efficiency.
Research presents enhanced Membership Inference Attacks (MIAs) on diffusion models using a frequency-domain approach, improving data privacy breach detection.
Research explores compiling discrete programming structures, like conditionals and iteration, into differentiable recurrent neural networks.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion