RESEARCHMonitorNEXT 12 MONTHS
Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Researchers introduced AgentCIBench, an evaluation harness measuring privacy risks when computer-use agents leak context across applications.
OneBench interpretation
Institutional assessment
So what
Cross-application agents risk leaking sensitive information between contexts, exposing institutions to unintended data exposure.
Do what
Review internal agent evaluation frameworks with the team responsible for AI risk management.