Small edits, large models: How Wikipedia advocacy shapes LLM values
Research shows a small group of Wikipedia editors can shape LLM values related to animal welfare, due to Wikipedia's dataset weight.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research shows a small group of Wikipedia editors can shape LLM values related to animal welfare, due to Wikipedia's dataset weight.
Research identifies a critical bug in common LLM repetition penalty implementations across inference engines like HuggingFace and vLLM.
Research identifies specific linguistic features in fine-tuning data that shift Llama-3.2-1B's reasoning preference towards pro-animal-welfare stances.
Researchers introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model, claiming trillion-parameter performance via agent-horizon scaling.
Research explores if Speech Language Models (SLMs) exhibit human-like sound symbolism, mapping speech acoustics to perceptual qualities.
DeepSearch-Evolve introduces a self-distillation framework for training web agents within DeepSearch-World, a verifiable environment.
Researchers propose an interpretable network framework using conceptual features to represent idiomatic and figurative meaning across eight languages.
Research finds longer prompts increase users' psychological ownership of AI-generated content; specific UI techniques can encourage longer prompts.
Researchers introduced PC-Mix, the first dataset for detecting partial-component audio spoofing in mixed speech and environmental sound conditions.
GigaChat Audio introduces a time-aware audio LLM that explicitly grounds answers with timestamps over up to 120 minutes of audio.
BizFinBench.v2 is a new bilingual LLM benchmark for finance, using real user query-response data for offline and online evaluation.
Research explores workload-driven optimization for on-device real-time English-to-Traditional-Chinese subtitle translation, focusing on low-latency, privacy, and short-context challenges.
Research proposes a 'Tool-Adaptive LLM Reranker' to reduce hallucination and latency in LLM-powered information retrieval by selectively invoking external tools.
Research demonstrates an LLM-based system for forecasting merger arbitrage outcomes by reasoning over hundreds of pages of technical M&A documents.
Research proposes an abstention-aware reinforcement learning method to mitigate LLM search-agent hallucinations by penalizing fabricated answers.
Research details adapting an open-source spoken language model for multilingual Singaporean contexts, demonstrating efficient fine-tuning methods.
New research proposes PTEI, a framework to integrate personality traits into LLMs to enhance emotional intelligence, addressing current underperformance.
GRADE proposes a hierarchical multi-agent system using learned gates to manage agent selection, hierarchy depth, and inter-agent communication efficiently.
Research introduces a human-AI triage model for POS fraud detection in Nigeria to combat algorithmic bias from infrastructure noise.
Researchers propose PolyInterview, an LLM-based platform for mock job interviews offering adaptive dialogue and multimodal assessment.
Anamnesis is an open-source platform for large-scale, demographically controllable survey simulation using LLMs and narrative backstories.
Research proposes Progressive Tree Drafting, a speculative decoding method to unlock parallelism in autoregressive LLMs, reducing inference overhead.
The first ChineseBabyLM challenge aims to train data-efficient Chinese LMs from scratch using 100M tokens, evaluated on NLU, cognitive alignment, and Hanzi knowledge.
Research identifies two sources of instability in LLM-based stance analysis: upstream data preprocessing pipelines and downstream LLM annotation.
AgentCheck is an open-source web workbench designed for reproducing, intervening, and mitigating failures in LLM agents that use tools.
Research explores pipeline optimizations for Natural Language to SQL (NL2SQL) translation, integrating NatSQL, synthetic data, and a reranker.
A study found LLMs exhibit 'epistemic paternalism' and differential refusal when tutoring marginalized students on sensitive historical topics.
HyperSafe proposes an inference-time method to restore safety alignment in fine-tuned LLMs without retraining or impacting task performance.
New research introduces MJ, a method for multi-turn LLM jailbreaking using decomposed credit assignment, identifying individual turn contributions.
Research explores using LLMs to reproduce human behavioral biases in route choice, aiming for scalable behavioral modeling with CPT parameters.
HiQA introduces a hierarchical contextual augmentation for RAG to improve multi-document QA, reducing hallucinations in language models.
Research paper introduces PM-KVQ, a progressive mixed-precision KV Cache quantization method to reduce memory overhead for long Chain-of-Thought (CoT) LLMs.
New arXiv research proposes DEER, a benchmark to evaluate deep research agents in generating expert-level reports, addressing current evaluation challenges.
RegCheck is a proposed AI tool designed to automate structured comparisons between research study registrations and their final published papers.
A new open-source Python package, GRADIEND, is released for gradient-based feature learning in language models, enabling persistent model rewriting.
Research presents PhoneticXEUS, a model for universal phone recognition trained on large-scale data, addressing multilingual generalization issues.
ProgramTab, a new research approach, proposes using a programmatic paradigm to enhance LLM reasoning over large and complex tabular data.
TechCrunch notes that established tech companies are intensely pursuing AI, driven by fear of missing out and profit potential.
Codex usage reportedly increased over 10x in six months, reaching 7 million users, with 1 million new users in approximately one day.
Uber's CPO discusses financial services ambitions, Waymo relationship, AV Labs data, and AI's role in rider/driver experience.
Apple ML Research introduces PARE, a framework for building and evaluating proactive AI agents by simulating stateful, sequential user interactions.
Apple Music developed a 305M-parameter multilingual semantic retrieval system, fine-tuned from GTE-multilingual-base, for cross-lingual search.
Video-generation startup PixVerse secured $439M in funding, pushing its valuation beyond $2B to expand its world model offering.
Microsoft Threat Intelligence identified ShinyHunters threat actor activity targeting SaaS applications via OAuth abuse, vishing, and supply-chain compromise.
Apple has revamped Siri, making it a core component of the iPhone user experience, available in the iOS 27 public beta.
OpenAI's new GPT-5.6 Sol, Terra, and Luna models are now generally available on Amazon Bedrock, expanding cloud-hosted frontier model access.
Palantir co-founder Joe Lonsdale's firm, 8VC, closed a $1.5 billion fund, focusing on defense tech and AI startups amid an investment boom.
8VC founder Joe Lonsdale claims China actively appropriates US AI and life-science IP, suggesting companies forgo patents to avoid theft.
Apple's IP lawsuit against OpenAI threatens to derail OpenAI's device strategy, according to Bloomberg, prior to case resolution.
An AI-fueled stock rout in South Korea impacted US markets; Apple filed a trade secrets lawsuit against OpenAI; FCC discusses space regulation.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion