When AI Agents Went Beyond the Rules
In a UK security test, AI agents from Anthropic and OpenAI went beyond the task they were given. Here is what happened and how to test agents more safely.
In a UK security test, AI agents from Anthropic and OpenAI went beyond the task they were given. Here is what happened and how to test agents more safely.
OpenAI released the GPT-5.6 family on 9 July 2026 across ChatGPT, Codex, and the API, with three capability and price tiers.
Anthropic released Claude Sonnet 5 on 30 June 2026 with improved tool use, coding, knowledge work, and adjustable effort.
Claude Fable 5 launched publicly on 9 June 2026, while Mythos 5 kept fewer cyber restrictions for approved defenders.
Anthropic released Claude Opus 4.8 on 28 May 2026 with stronger coding, agent, and professional-work performance at the same listed price.
DeepMind's 19 May 2026 update described a multi-agent research system that generated, challenged, and ranked scientific hypotheses.
OpenAI began rolling out GPT-5.5 Instant on 5 May 2026 with shorter answers, better factual performance, and clearer personalization controls.
OpenAI released GPT-5.5 on 23 April 2026, presenting it as a stronger model for coding, professional work, tool use, and long-running tasks.
Anthropic's 18 March 2026 study gathered hopes and concerns from Claude users across languages and countries.