RESEARCHInvestigateNEXT 12 MONTHS
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Researchers released BAD-ACTS, a benchmark and taxonomy for evaluating LLM agent robustness against adversarial attacks and harmful actions.
Open source