Catalyst-Agent: Autonomous heterogeneous catalyst screening with an LLM Agent
Research paper proposes 'Catalyst-Agent,' an LLM agent for autonomous screening of heterogeneous catalysts using MLIPs and GNNs.
Search signals, briefings, company results, benchmarks and glossary terms.
Search signals, briefings, company results, benchmarks and glossary terms.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Research paper proposes 'Catalyst-Agent,' an LLM agent for autonomous screening of heterogeneous catalysts using MLIPs and GNNs.
Research explores scaling zero-RL, a method using verifiable rewards without human data for chain-of-thought reasoning, to trillion-parameter models.
Research introduces a policy-conditioned constrained decoding method for Text-to-SQL systems, enhancing column-level access control.
Research introduces a paired analysis methodology for LLM safety, tracking risk from prompt to response across harm categories and severity.
Research paper argues existing LLM evaluation metrics for dialogue (BLEU, ROUGE) fail to capture deeper conversational quality like coherence and consistency.
Research introduces an 'Attention-Discounted Adaptive Sampler' for masked diffusion language models to improve parallel token generation in inference.
Research explores methods to scale 'point-in-time' LLMs, eliminating future data leakage for valid financial backtesting and causal inference.
Research introduces MAGE, a framework to analyze stability-performance trade-offs in multi-component prompt optimization for LLMs, focusing on component interaction.
Research identifies internal mechanisms in LLMs that separate a character's beliefs from reality, involving a value slot and a query router.
Fin-Analyst, a hybrid LLM trading agent with eight specialist LLM pipelines and rule-based signals, is proposed for FinMMEval 2026 Task 3.
Research explores interactive multi-feature fusion to improve continuous semantic reconstruction from non-invasive brain recordings, moving beyond static representations.
CityBehavEx is a new LLM-assisted urban simulation platform that claims to scale to city-size populations with empirically validated mobility patterns.
Research finds that context-reduction layers for coding agents do not consistently lower actual billed costs, despite reducing text volume.
Research models how speakers of related languages achieve partial intelligibility (intercomprehension) using a Bayesian framework and L1 language models.
Research identifies "pigeonholing," where unintentionally bad prompts degrade LLM performance and cause mode collapse, even without malicious intent.
PerfCodeBench is a new benchmark for evaluating LLMs on system-level high-performance code optimization, focusing on efficiency over correctness.
Researchers propose a fine-tuned multi-agent framework leveraging LLMs to detect OCEAN personality traits more accurately from long narratives.
New research introduces the Roundtable Context Window Test (RCWT) to measure the "task-budget displacement" from coordination content in LLM prompts.
FinResearchBench II introduces a scalable pipeline for evaluating deep research agents by automatically synthesizing financial report quality rubrics.
Researchers identified a 'countdown subcircuit' in Llama 3.1-70B-Instruct enabling precise token-tracking across various tasks.
Research argues current AI alignment practices, focused on measurable optimization, overemphasize fluent outputs without addressing deeper value questions.
New research introduces MemOps, a benchmark evaluating the lifecycle of memory operations in LLM-based agents across long-horizon conversations.
Research explores continual learning for medical visual question answering, focusing on model adaptation to new tasks without forgetting past knowledge.
GRID (Grammar-Railed Decoding) is a new method for generating SQL that guarantees syntactic validity, policy adherence, and compliance-grade records.
LakeQuest is a new research benchmark for evaluating question answering systems over heterogeneous, weakly structured data lakes.
Research introduces 'Speculate with Memory,' a method to accelerate LLM agents using smaller, cheaper models with learned memory systems.
Research paper Code-MUE proposes a method for measuring the uncertainty of Code LLMs through execution-based semantic interaction graphs.
Research introduces ATLAS, an adaptive testing framework using Item Response Theory for LLM evaluation, aiming to reduce cost and time.
Research proposes a 'function-aware fill-in-the-middle' pretraining method for coding agent foundation models to improve tool integration.
Research suggests current LLM agents often over-contextualize simple tasks, lacking awareness to estimate effort, leading to inefficient execution.
Research finds that Group Relative Policy Optimization (GRPO) for small language and vision-language models mostly reshapes existing behavior, rather than adding new skills, particularly under certain learning rate conditions.
New research proposes methods for attributing failures in LLM-based agentic systems, moving beyond costly prompting or manual annotation.
A new framework, G-SHARE, uses guideline-based structured reasoning for human-factor event diagnosis, aiming to improve consistency over LLMs.
Research finds state-of-the-art LLMs perform poorly on bidirectional Braille translation, revealing significant accessibility failures.
Research from arXiv explores using LLMs to model consumer story expectations by generating continuations and extracting interpretable features.
Research proposes a graph-based framework combining weak supervision and propagation analysis to detect disinformation narrative diffusion on Telegram channels.
Research outlines ontology-amplified distillation of a Qwen3.6-27B model for sovereign enterprise LLM deployment, addressing data residency needs.
Research explores LLMs for learning chemical reaction mechanisms, aligning stepwise deduction with LLM reasoning paradigms beyond coarse-grained reactions.
FairCoder, a new benchmark, tests LLM social bias in high-stakes decision-making by framing tasks as code generation to reveal implicit bias.
OpenAI researcher Miles Wang is reportedly in talks to launch an AI drug discovery startup, seeking a $2 billion valuation.
Apple ML Research proposes a method to adapt pretrained visual encoders for image generation with a single latent layer, improving efficiency and quality.
Apple ML Research proposes CLaRa, a unified RAG framework using embedding-based compression and joint optimization to reduce context length.
Apple ML Research proposes a framework for quantifying uncertainty in LLM function-calling to mitigate risks in autonomous task execution.
Hugging Face introduces VoiceEQ, a new metric for evaluating the human quality and emotional expressiveness of voice AI systems.
Codex, a programming assistant, is reportedly gaining 1 million new users daily, indicating significant adoption of AI coding tools.
Expert commentary from AINews highlights 'agentic AI' and a shift to building AI systems around agents from the AIE World’s Fair 2026.
Musician Lorde commented on stage that AI glasses are 'not sexy' and expressed concerns about distinguishing reality.
The Financial Times reports on market's intolerance for perceived failures, citing IBM as an example.
OpenAI's GPT-5.6 Sol model is reportedly deleting user files and data; OpenAI acknowledged a related issue in June.
OpenAI is reportedly planning to announce a screenless ChatGPT smart speaker this year, featuring a camera and sensors for environmental understanding.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion