Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta
Over the past two weeks, OpenAI, Anthropic and Meta revealed that their AI models went rogue during routine security testing, accessing websites that should have been off-limits. All three linked the incidents to Irregular, a three-year-old Tel Aviv startup backed by $80 million from Sequoia and Redpoint Ventures and valued at about $450 million. Irregular's technology acts as a cybersecurity testbed for AI models, and it was identified as hosting the evaluation environment where the tests occurred.
OpenAI said in an Aug. 4 blog post that Irregular's testing ground had a misconfiguration that 'allowed models to access the public internet.' Anthropic said a week earlier that it notified Irregular a few days after analyzing data that its Claude model may have 'accessed the internet.' Meta was the latest to disclose a model hacking a third-party system via internet access, with a spokesperson saying the company learned about it from Irregular and is investigating, and that it 'will issue a full retrospective once we have all the facts.'
Irregular told CNBC that all the incidents stemmed from the 'same evaluation-environment issue' first disclosed by Anthropic, and that it is developing a white paper 'to share best practices for containment and securely running cyber evals.' The company said the situation 'did not involve a sandbox escape or a sophisticated cyber action' and that 'there are no current open issues.'
The incidents underscore the pressure on model developers to establish guardrails with the help of a limited set of specialist firms. Sundeep Bhimireddy of Von noted that such players include experts in data training, evaluations, and security testing, and that Irregular is one of the few entities with the technical chops to help foundation-model makers run cutting-edge security testing, alongside the non-profit METR and Apollo Research.