RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
RefusalGuard introduces a geometry-preserving fine-tuning method to prevent downstream task tuning from degrading LLM safety refusal behavior.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
RefusalGuard introduces a geometry-preserving fine-tuning method to prevent downstream task tuning from degrading LLM safety refusal behavior.
Research introduces RouteScan to audit MoE LLM safety using expert routing telemetry without exposing prompt or output content.
New research shows code LLMs can leak proprietary training logic even when syntax and structure differ substantially from source code.
Researchers introduced Mint-Agent, a family of finance-native agentic models designed for auditable long-horizon financial research.
Researchers propose 'The Metanym Game', a self-contained, contamination-resistant LLM benchmark evaluating structural intelligence.
Research introduces XIH-Bench, a new benchmark to evaluate instruction hierarchy compliance in multilingual LLMs, finding language impacts adherence.
Researchers propose Zhijing, a framework for measuring and integrating social intelligence in LLMs, including the SoMBench benchmark.
Research overviews Bayesian and frequentist simulation-based inference with machine learning for inverse problems and parameter estimation.
Bilibili's Index-1.9B is a new series of 1.9B parameter open small language models, pre-trained on 2.8T predominantly Chinese and English tokens.
Research introduces Calibrate-Then-Delegate (CTD), a model-cascade approach for LLM safety monitoring that uses a cheaper model to screen and delegates hard cases to an expert, optimizing for cost and accuracy.
FT reports Anthropic enterprise customers are shifting workloads from flagship models to cheaper options to manage inference costs.
The London Stock Exchange plans to launch LSE 24 in H1 2027, a 24/5 trading venue targeting algorithmic and agentic ETP trading.
Anthropic's flagship Fable 5 model sees sluggish corporate adoption as enterprise buyers opt for cheaper alternative tools.
Nvidia server prices are rising over 15% due to soaring memory chip costs, impacting large infrastructure deployments.
Apollo's Chief Economist Torsten Slok reports AI adoption is suppressing wage growth before causing direct job cuts in labor markets.
Latent Space analyzes how frontier models are absorbing external orchestration harnesses directly into their underlying weights.
DBS announces the deployment of agentic AI capabilities for corporate bankers to automate routine tasks and enhance client engagement.
Coupa demonstrated a Payment Batch Creation Agent processing 2,395 payments across 14 batches with mandatory human sign-off.
Nvidia research demonstrates agentic harness engineering and task fine-tuning can achieve high performance on smaller base models.
Nvidia licenses Poolside's Model Factory software for $6 billion and plans to acquire 109 employees to bolster model development.
AWS introduced ADOP, a Bedrock reference architecture using agents to automate Bronze-to-Gold data pipeline lifecycles.
AWS introduced Amazon Bedrock AgentCore Gateway, a framework to govern and audit AI agent tool access across enterprise infrastructure.
AWS details a query-aware context compression pattern on Bedrock using a smaller model to filter retrieved chunks and lower RAG costs.
NVIDIA details AdaptGrow, a GPU-accelerated algorithm for converting rolling correlation and tail-dependence matrices into factor clusters.
Dutch regulators hit Uber with an €825m fine for using automated systems to suspend driver accounts without sufficient explanation.
NVIDIA introduced AVO, an agent harness architecture that achieves high benchmark performance on long-horizon reasoning tasks.
NVIDIA outlines architectural principles for integrating security and trust controls into long-horizon AI agent execution stacks.
DBS is deploying agentic AI to 1,500 bankers globally to automate corporate credit assessment workflows.
Researchers introduced VSysBench, a new benchmark evaluating how multimodal LLMs adhere to system-message constraints alongside visual inputs.
Researchers introduced ReCache, a framework enabling KV cache reuse across modular agent tool schemas to reduce inference latency.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion