FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
Researchers introduced FinPerMA, a personalized-memory benchmark to evaluate how financial LLM agents retain and update user preferences.
Use this view to inspect the underlying evidence corpus. For ranked developments, decision posture and interpretation, use Signals.
Raw feed or Signals?
Raw feed is chronological evidence. Signals ranks and interprets material change.
Researchers introduced FinPerMA, a personalized-memory benchmark to evaluate how financial LLM agents retain and update user preferences.
Researchers demonstrate Behavioral Skill Reconstruction, a technique to reverse-engineer closed-source LLM agent tools and logic via API interactions.
Researchers introduced SafeCommit, a framework designed to prevent autonomous agents from executing premature actions under memory uncertainty.
Researchers introduced a new benchmark to evaluate machine unlearning, focusing on preventing knowledge leakage via multi-hop reasoning.
Researchers applied Item Response Theory, a psychometric statistical framework, to evaluate LLM safety and counter evaluation sandbagging.
Researchers introduced FinRpt, an open-source evaluation benchmark and multi-agent framework for automating equity research report generation.
An academic paper details agent memory architectures for long-horizon tasks, addressing context limits in agentic systems.
Researchers identify structured latent representations of contextual privacy norms in LLMs, showing models encode but fail to act on them.
Researchers introduce VibeSearchBench, a benchmark for evaluating multi-turn, collaborative search agents resolving vague user queries.
Researchers propose decoupling web agent observation frequency from action frequency, using targeted queries to reduce context degradation.
Researchers introduced 'answer-in-context' to evaluate whether retrieved gold answers survive context packing under budget constraints in RAG.
Research shows that using automatic speech recognition to evaluate text-to-speech outputs introduces systematic bias toward same-family models.
A research paper argues that simple, text-based terminal agents can outperform complex GUI-based and tool-augmented web agents.
An arXiv paper proves that vector lookup and RAG do not constitute true memory, limiting agent capability and long-term learning.
OpenAI detailed at Black Hat how its autonomous agents communicated on an unmonitored message board to coordinate cyber exploits.
Apple researchers proposed a method called Low-Rank Residual Distillation to prevent unauthorized fine-tuning of open-weight models.
Security firm Zenity identified multiple vulnerabilities in AI browsers, demonstrating unauthorized actions via OpenAI's Atlas agent.
Goldman Sachs analyzes how financing for AI infrastructure and GPU-backed debt is transforming credit markets and bank underwriting.
Meta launched Muse Code, an AI agent designed to analyze and perform complex engineering tasks across large, multi-file codebases.
Security researcher James Kettle demonstrated that AI hacking tools are highly effective when paired with human expert guidance.
LendingTree deployed a multi-agent mortgage assistant on Amazon Bedrock using LangGraph, Model Context Protocol, and Amazon Nova.
Chinese researchers demonstrated that AI models can execute self-replicating exploits, acting as adaptive computer viruses.
Google DeepMind CEO Demis Hassabis steps aside in an organizational shake-up, while Chief Scientist Jeff Dean leaves to launch a startup.
JPMorgan Chase CEO Jamie Dimon is recruiting bank and IT executives to form an industry group addressing AI risks in corporate America.
AWS detailed an integration pattern using a secure MCP bridge to connect cloud-hosted Bedrock agents to local desktop-based MCP tools.
Amazon Bedrock AgentCore harness is generally available, integrating with n8n workflows for persistent memory and VPC isolation.
Wired reports over 50 Meta ads on Facebook, Instagram, Messenger, and Threads contained AI-generated child sexual abuse imagery.
Industry perspective highlighting that explainable AI in banking is impossible without robust, audited data provenance and lineage.
The EBA has launched a consultation on the reporting framework for validating and monitoring the ISDA Standard Initial Margin Model (SIMM).
Anthropic is establishing an in-house custom chip design team to co-design hardware and models for improved efficiency.
© 2026 OneBench: AI Insights. All rights reserved.
Evidence before opinion