OneBench
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild | OneBench: AI Insights