Most published safety evaluations for language models are single-turn or short-conversation, text-in/text-out. Most deployed risk now comes from agents that call tools, take multi-step actions, and operate over long horizons with real side effects (sending an email, running a shell command...