RESEARCHMonitorNEXT 12 MONTHS
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Research across 62,800 evaluations shows agentic scaffolds like ReAct and multi-agent loops significantly alter measured model safety.
Inspect the evidence
- Inclusion basis
- Enterprise AI
- Publisher and source type
- arXiv cs.CL — Computation and Language · RESEARCH
- Published by source
- 25 September 2026
- Collected by OneBench
- 26 Sept 2026, 03:01 UK
Stored source excerpt
arXiv:2603.10044v3 Announce Type: replace-cross Abstract: Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex…
Short excerpt from the collected text, not the full source. Use the source link to read it in context.
The factual summary is a OneBench synthesis, not a quotation or independent verification. Collection time is not publication time. Open the source for its full context; related reporting can share the same underlying announcement.