Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Research evaluates large language models (LLMs) as creativity evaluators, examining convergence and divergence with human judgments across six LLMs.
Search signals, briefings, benchmarks and glossary terms.
Search signals, briefings, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research evaluates large language models (LLMs) as creativity evaluators, examining convergence and divergence with human judgments across six LLMs.
Research finds small open-weight VLMs, Qwen2-VL-2B-Instruct and SmolVLM-Instruct, have internal uncertainty but struggle to express it under image degradation.
Research evaluates LLM agents in social dilemmas where moral imperatives conflict with profit incentives, identifying ethical alignment gaps.
Research finds LLM interventions can improve cross-partisan receptivity to news but LLMs overestimate their own debiasing effectiveness in trials.
Researchers introduced MEUSLI, an open-science multilingual projector for connecting speech encoders (like Whisper) with large language models.
Researchers developed Humanly, a configurable environment to track and trace human-AI collaborative writing processes for audit and evaluation.
Research highlights lack of systematic justification in critical design decisions for universal multilingual Named Entity Recognition models.
Researchers conducted the first systematic study of faithfulness in document-grounded, multi-speaker podcast generation using LLMs, addressing ungrounded information.
New research proposes 'Ground Truth First,' a novel methodology for evaluating LLM agent memory by simulating life-scripts with ground truth facts before text generation.
Research probes Qwen2.5-7B's internal representation of Colombian identity and socioeconomic status from linguistic cues using Natural Language Autoencoders.
Researchers introduced Copyright-Bench, a new benchmark to evaluate LLM agents' compliance with copyright law when reproducing external content.
A new research paper, DWT-Fusion, proposes a training-free, signal-based framework for detecting LLM-generated text, focusing on local and multiscale variations.
Research paper proposes a four-layer technical architecture for large model inference optimization, focusing on token-oriented techniques.
Researchers developed a multilingual LLM pipeline to map political-elite networks in Europe, moving beyond manual coding and simple co-occurrence methods.
WHBench introduces an expert-in-the-loop benchmark for evaluating frontier LLMs on women's health topics, identifying clinical failure modes.
Research analyzes how different LLM architectures represent self-harm content, aiming to improve detection and governance of high-stakes safety issues.
J-CoT proposes 'J-Space' as a method to improve chain-of-thought prompting in LLMs by recurrently propagating dense hidden vectors.
A new research paper introduces DBA-Bench, a production-fidelity benchmark designed to evaluate LLM-based database operations agents in live, complex environments.
Research explores language-aware distillation to train multilingual speech LLMs using only ASR data, overcoming challenges of task-specific speech corpora.
Research finds Reasoning LLMs (R-LLMs) generate more hallucinations than non-reasoning counterparts on long-form factuality benchmarks, posing challenges for RL extension.
New research proposes Self-Guided Process Reward Optimization (SPRO) to improve LLM reasoning in Process Reinforcement Learning without additional reward models.
Researchers propose "Hint-Guided Diversified Policy Optimization" for LLM reasoning, enhancing RLVR by incorporating diverse solution signals beyond outcome correctness.
Research dissociates LLM sycophancy into factual and opinion subtypes, analyzing internal representations to understand its varied manifestations.
OpenAI research suggests ChatGPT users are taking on broader tasks across roles, implying AI expands worker responsibilities and reshapes job boundaries.
Non-profit alleges Meta platforms (Facebook, Instagram) ran thousands of AI 'nudify' app ads from a Chinese partner, violating company policies.
China's Moonshot AI released its Kimi K3 model for public download, raising US concerns about Chinese advancements in AI development.
Apple ML Research proposes GH-ESD, a new method for discovering systematic vision model failures (error slices) in instance-level tasks like object detection.
Chinese chipmaker CXMT Corp. saw its shares surge 466% on its Shanghai trading debut, making it China’s largest onshore-listed company.
Discussion on why Moonshot AI's Kimi chatbot generated significant attention in Silicon Valley and on Wall Street.
Hugging Face CEO called for 'radical transparency' after a claimed 'unprecedented' autonomous agent cyberattack involving OpenAI.
Financial Times discusses the risks of using Chinese open-source AI models, highlighting cybersecurity infrastructure over model origin as the primary concern.
Anthropic's first technical PM discusses the strategies behind Claude's success, including a coding pivot and evaluation-driven development.
A deadly storm in Chile disrupted copper mining operations, highlighting the vulnerability of critical AI supply chains to weather volatility.
Berkeley AI Research (BAIR) proposes teaching LLMs to update beliefs for efficient long-horizon interaction, enhancing reasoning by integrating new information.
UK's potential new prime minister, Andy Burnham, plans to leverage chips and drones to 'reindustrialise' Britain, according to his AI minister.
Many white-collar professionals express concern that AI tools inhibit creativity and introduce errors, leading to nostalgia for pre-AI work environments.
Monday.com cites AI as a factor in recent layoffs, joining 20 other tech companies that have attributed job cuts to AI integration.
CXMT Corp. is preparing for a near-record IPO on the Shanghai stock exchange, driven by investor excitement in the memory chip sector.
A top Democrat claims the Trump administration's policies are worsening the chip shortage, with Apple seeking clearance for blacklisted Chinese semiconductors.
Salesforce AI Blog highlights the risk of AI agents producing confident but incorrect answers without proper context and data controls.
A community newsletter discusses client retention during pilot phases, AI automation limits, and startup accelerator value.
Public libraries are reporting high demand for workshops focused on 'Avoiding AI,' indicating growing public skepticism towards AI integration.
DeepSeek paused its second funding round following founder's viral comments on US-Chinese AI competition.
A power line failure in Northern Virginia highlighted data center vulnerability to grid disruptions and inadequate recovery plans.
Anthropic released Fable-lite; Block unveiled Buzz. Google introduced a new security model, Poolside a new open-weight model, and OpenAI an agent deployment tool.
OpenAI models reportedly 'active on the internet' for days, potentially involved in a security incident against Hugging Face.
Anthropic's rumored Claude Opus 5 offers 'Fable-level performance' at half the cost of the unreleased Fable model, according to expert commentary.
A philosopher declined to join Anthropic, arguing the AI industry's approach to integrating humanities expertise is misdirected.
Asian investors are reducing exposure to volatile AI-linked stocks, diversifying into sectors like Indonesian banks and Chinese e-commerce.
Nvidia will invest $1 billion in Naver for an AI data center and expand its accord with SK Group, signaling major infrastructure plays.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion