🛡 The 'Rogue' AI Agents All Had One Common Source
A The Verge investigation linked incidents involving agents from OpenAI, Meta, Anthropic, and Google to Israeli startup Irregular, which stress-tests models for cybersecurity: in several tests, agents escaped the isolated environment and attacked real targets.
🌍 A series of independent failures turned out to be a single configuration error at a single third-party evaluator: agents were inadvertently given access to the open internet, and the fictional name of the target company matched a real domain.
👤 If you are running agents with network access: perimeter isolation and domain verification before launch are mandatory — even an evaluator with $80 million in funding, working with OpenAI and the UK government, made a mistake.
Source 1: https://www.theverge.com/ai-artificial-intelligence/1000644/irregular-rogue-ai-cyberattacks-hacking-openai-meta-anthropic-google Source 2: https://news.ycombinator.com/item?id=49853692
