Position: Natural Language Should Not Fully Replace Formal Languages
Research argues natural language should not fully replace formal languages for software design due to its inherent underspecification.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research argues natural language should not fully replace formal languages for software design due to its inherent underspecification.
A retrieval-augmented, multi-agent LLM framework with human-in-the-loop improved accuracy and reduced review time for medical notes.
UZH Shared Task 2026 winner LLM-INSTRUCT uses open-weight 8B parameter models for paragraph-level argument mining with constrained structured prediction.
AlphaAgent is a new skill-driven agent framework that separates retrieval from generation for complex literature analysis in materials science.
Research identifies epanorthosis, a rhetorical self-correction, as a systematic overuse in LLM text due to training data and RLHF.
Research explores expert-aware contrastive decoding in Mixture-of-Experts (MoE) models to mitigate LLM hallucinations, extending prior work on transformers.
Research claims Mixture-of-Experts (MoE) routing in models like Phi-3.5-MoE and Gemma-4-27B-A4B aligns with Huffman Coding principles, uncovering a Frequency-Diversity Law.
Researchers propose a multi-axis evaluation framework for structured audio captions, addressing limitations of existing metrics for multimodal outputs.
New research, "Moir," proposes a knowledge editing method for LLMs that mitigates degradation of reasoning capabilities by letting the model direct its own edits.
Research explores methods like low-rank decomposition and quantization to compress large language models, aiming to mitigate performance degradation at high ratios.
Research introduces TopoGuard, a graph theory-based defense against split-knowledge attacks on RAG systems where individually benign documents create false associations.
Research identifies 'Routing Subspaces' to locate why fine-tuned LLMs may appear safe in evaluation but still exhibit problematic behaviors in use.
Research explores methods to make watermarks in open-source LLMs durable against post-training modifications like model merging.
Research identifies "directional hallucinations" and ideological drift in LLMs when answering political questions, using a new measurement framework.
Telco-GAIA introduces a bilingual, multi-modal benchmark for evaluating tool-using agents on real-world telecom data with complex reasoning.
Research investigates the effectiveness of LLMs in detecting their own generated content across programming and writing tasks.
GenDB, a generative query engine using LLM agents, is demonstrated to automatically generate customized query processing code, aiming to reduce engineering effort.
PersonaTrail introduces a new benchmark for personalized web agents, evaluating their ability to infer context from user browsing histories.
New research proposes "Answer-then-Edit" for anti-distillation, generating defensive LLM outputs to prevent unauthorized knowledge extraction while preserving utility.
Research explores factors impacting the detection of deceptive outputs from LLMs, noting current probes fail in out-of-domain scenarios.
Research introduces WaveformQA, a new benchmark to evaluate LLMs' temporal reasoning over digital waveform data, addressing a design verification gap.
Research finds language models widely hallucinate in chemical reasoning, often producing correct answers with fabricated intermediate steps.
Research explores using transformer-assisted LLMs for source code summarisation to improve secure software development lifecycle maintenance.
NVIDIA-labs introduces Object-Oriented Agents (NOOA), a Python framework for building reliable AI agents by representing agents as Python objects.
Researchers propose "Refusal-Gated Decoding" to maintain LLM refusal behaviors when using high-temperature sampling for output diversity.
Research finds LLMs can generate deceptive responses with high confidence, increasing their persuasiveness to users, across various models and datasets.
Research finds all frontier LLMs exhibit response drift, producing outputs that deviate from expert-validated references, uncharacterised by human evaluation.
Euclid-MCP proposes a standardized protocol server for integrating LLMs with Prolog-based symbolic reasoning, aiming for reliable logical outputs.
VibeVoice-ASR-BitNet introduces a highly compressed ASR model using INT8 and BitNet-style ternary quantization for edge CPU real-time inference.
Research paper introduces Semantic Field Theory (SFT) as a computational model for lexical semantics and stabilized interpretation, refining its mathematical core.
Research benchmarked five LLMs on multi-sensor physical hazard assessment, finding all consistently failed to issue precautionary warnings.
MedGame introduces an LLM-powered framework to transform static clinical cases into interactive, decision-centered storytelling games for medical education.
Research introduces HiMe, a real-time, self-hosted, open-source personal agent platform for health insights from wearable data using LLM agents.
Research explores how inherent narrative structures and archetypal roles in LLM training data systematically influence model behavior and pose governance risks.
Instruct-FD introduces an instruction-conditioned benchmark for evaluating controllable turn-taking in full-duplex spoken dialogue systems.
New research proposes THOR, a Theta-Gamma hierarchical oscillatory reasoning framework, to improve multi-hop question answering by addressing attention decay and error accumulation.
A preliminary research study explores whether valence (emotional tone) in natural language can reflect morality, impacting AI ethics.
CAMeR introduces a keyword-gated hybrid activation memory framework for LLM agents to selectively retain relevant information over extended dialogues.
Pulsar Attention proposes a new method for distributed LLM inference that replaces static context anchors with content-aware components, reducing compute costs.
TextGrad, a method for optimizing language model text components from natural language feedback, struggles with agentic systems due to delayed feedback attribution.
Research tested GPT-4.1's ability to predict opinions using persona simulation, accurately forecasting 2024 election outcomes based on U.S. state personas.
Research introduces REGARD to study affective framing differences in LLMs, moving beyond simple sentiment to understand regional biases.
REFACT proposes an adaptive method for LLMs to restate facts in their chain-of-thought reasoning, aiming for compact, faithful, and context-aligned outputs.
Researchers introduced Rushes, a human preference dataset for pluralistic alignment collected from interactive AI-generated narratives.
Research fine-tunes small language models (0.6B-20B parameters) to translate natural language into MiniZinc, a domain-specific constraint language.
Tencent introduces WorkBuddy Bench, a multi-domain coding-agent evaluation suite with contamination-resistant task construction methodology.
LegalCiteTrust benchmark evaluates citation trustworthiness in LLM-generated long-form legal research by identifying misrepresentations.
Research proposes using reinforcement learning to detect user interface (UI) principle violations, including accessibility and poor visual hierarchy, in LLM-generated front-end code.
Researchers propose Constrained Shared-Private Fusion (CSPF), a new method for more reliably evaluating non-verifiable AI tasks by integrating diverse reward models.
RUMBA introduces a new Russian-language benchmark for long-term conversational memory in LLMs, focusing on granular retrieval and reasoning.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion