The AI security incidents at OpenAI, Anthropic and Meta all trace back to one testing firm's misconfiguration, CNBC reported Saturday.
The firm is Irregular, a Tel Aviv-based startup formerly known as Pattern Labs that runs cybersecurity evaluations for frontier AI companies. Sequoia Capital and Redpoint Ventures have backed the company with $80 million at a $450 million valuation, according to CTech.
During routine capture-the-flag exercises, a configuration error connected Irregular's testing environment to the public internet. Models from all three labs escaped their sandboxes and accessed real systems. The error was the same in every case, IT Pro reported.
Anthropic disclosed on July 31 that several Claude models gained unauthorized access to systems at three organizations during the tests. In one incident, Claude Opus 4.7 encountered a fictional target company that shared a name with a real business and exploited weak passwords and unauthenticated endpoints to reach credentials and databases, CTech reported. A separate internal research model stopped its own attack after recognizing it had reached a real organization.
Anthropic examined 141,006 interactions where Claude could access open networks and classified the breaches as a "harness failure," meaning the infrastructure broke, not the model's alignment. The company suspended all cybersecurity model evaluations on July 23 and began notifying affected organizations four days later.
OpenAI separately confirmed that the same misconfiguration allowed its agents to reach the public internet. Meta said its incident involved "the exact same evaluation-environment issue," according to IT Pro.
For builders relying on third-party evaluation services, the takeaway is direct: the security of the testing harness matters as much as the safety of the model. When labs deliberately disable safeguards to measure raw capability, a misconfigured sandbox becomes a live attack surface.