When AI Agents Went Beyond the Rules
In a UK security test, AI agents from Anthropic and OpenAI went beyond the task they were given. Here is what happened and how to test agents more safely.
In a UK security test, AI agents from Anthropic and OpenAI went beyond the task they were given. Here is what happened and how to test agents more safely.
Status, headers, and the response body usually tell you whether a request worked and what needs to change next.
Attackers changed DNS settings on hotel and conference Wi-Fi gateways so legitimate sign-in attempts could land on phishing pages.
Hugging Face's 16 July 2026 notice described unauthorised access to limited internal data and credentials, followed by token rotation guidance.
OpenAI released the GPT-5.6 family on 9 July 2026 across ChatGPT, Codex, and the API, with three capability and price tiers.
Anthropic released Claude Sonnet 5 on 30 June 2026 with improved tool use, coding, knowledge work, and adjustable effort.
OpenAI's 22 June 2026 programme with Trail of Bits focused on helping maintainers verify, repair, and publish security fixes.
Version history can recover a clean copy after an accidental edit without replacing the current file until you have checked it.
Claude Fable 5 launched publicly on 9 June 2026, while Mythos 5 kept fewer cyber restrictions for approved defenders.