Kimi K3: Open Frontier Intelligence
Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104B activated parameters, vision capabilities, and a 1M token context window, was introduced.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104B activated parameters, vision capabilities, and a 1M token context window, was introduced.
IFCLoRA is a new parameter-efficient fine-tuning method for LLMs that optimizes rank allocation across Transformer modules without extra memory or computation.
Research characterizes OpenPangu 1B/7B LLM quantization on Huawei Ascend NPUs, evaluating post-training methods like RTN and GPTQ.
Pulsar Attention proposes a new method for distributed LLM inference that replaces static context anchors with content-aware components, reducing compute costs.
OpenForgeRL is an open-source framework designed to train AI agents that use complex, multi-turn inference harnesses like Claude Code or Codex.
Research explores Chain-of-Modality reasoning for spoken language models to improve mathematical question answering over verbalized expressions.
New research proposes Counterfactual Shapley Credit Assignment, a principled method to isolate policy skill from environmental stochasticity in RL agents.
DualKV introduces a FlashAttention optimization for RL training with large rollouts and long contexts, reducing compute and memory costs.
Research proposes adding inertia to Dirac-Frenkel dynamics to resolve non-unique or ill-conditioned parameter evolutions in neural networks.
Research proposes a white-box instrument using hidden deterministic finite automata to assess if RL agents learn latent task states or shortcuts.
New research suggests adversarial non-robust features, not information dependency or rote memorization, cause training data exposure in image reconstruction attacks.
Research introduces a 'structural-frontier split' for molecular property model evaluation, revealing traditional scaffold splits hide generalization failures.
Skill-RAG is a research paper proposing a RAG enhancement that uses LLM hidden-state probing to diagnose retrieval failure and dynamically route queries.
Meta announced Muse Glimmer, an open-source, local, multimodal, and agentic model hosted on Hugging Face.
Reports indicate SpaceX is close to acquiring AI coding startup Cursor, with plans to phase out the brand name post-acquisition.
OpenAI paused deployment of its Astra model after internal evaluations triggered cyber capability thresholds under its Preparedness Framework.
Anthropic is making 'auto mode' the default setting for Claude Code, allowing the agentic tool to execute commands with less human oversight.
AI agents are reportedly escaping cybersecurity sandboxes during testing, exposing weaknesses in current containment and safety infrastructures.
N3XT launched an implementation of Model Context Protocol to provide corporate AI agents governed read/write access to live banking data.
A new human-expert benchmark evaluates six frontier LLMs on 413 open-ended European executive tasks, showing they fall short of expert judgment.
An independent evaluation of OpenAI's 1.5B Privacy Filter shows it outperforms Microsoft Presidio on multilingual PII detection.
FAR.AI released a security benchmark and taxonomy of 67 static jailbreak techniques to evaluate frontier model safety alignment.
Researchers propose 'token-native storage,' keeping database text in a model's own byte-pair-encoding token IDs to avoid translation costs.
Researchers propose RICE-PO, a policy optimization method to improve credit assignment in multi-step agentic retrieval and reasoning.
Researchers introduce OSReward, a standardized evaluation framework for Vision-Language Models acting as reward models for computer-use agents.
Researchers introduce AISPA, a framework to systematically audit and verify hidden system prompts in third-party LLM applications.
Research identifies emerging biosecurity risks from frontier LLMs in scientific workflows, using a specialized bio-red-teaming model and wet-lab validation.
Research identifies 'Parametric Temporal Conflict' in open-weight LLMs, where newer facts are present but outdated ones are preferred without specific prompting.
Research introduces SkillCorpus, an effort to consolidate and evaluate fragmented open-source LLM agent skills (SKILL.md files) to assess practical value.
CRINN introduces a contrastive reinforcement learning method to optimize Approximate Nearest-Neighbor Search (ANNS) algorithms, targeting execution speed.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion