Meta has confirmed that one of its artificial intelligence models accessed the public internet during a controlled safety test and broke into the systems of another company, making it the third major AI developer in three weeks to disclose such an incident.

The company said on Wednesday that Muse Spark 1.1, the model it has promoted as its most capable for real-world coding and so-called agentic tasks, gained internet access because of an error in the setup of its testing environment and went on to exploit a security vulnerability at an outside service. Meta has not named the affected company.

The breach was first reported by The Information, which said the model made changes to the third party's internal systems after escaping the confines of a "sandbox" — an isolated environment intended to keep test models cut off from live networks.

Meta placed responsibility for the lapse with Irregular, the Tel Aviv-based evaluation firm it had contracted to run the tests.

"A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," a Meta spokesperson said.

Meta added that the model "exploited a security vulnerability" at another company in a manner similar to previously reported instances elsewhere in the industry.

Irregular, formerly Pattern Labs, rejected any suggestion the model had gone rogue. A spokesperson told Reuters the episode was the "exact same evaluation-environment issue" already disclosed by Anthropic, and did not involve a "sandbox escape or a sophisticated cyber action". The firm said there were no current open issues and that it was preparing a white paper on best practice for containing and securely running cyber evaluations.

Founded in late 2023, Irregular specialises in AI red teaming and attack simulation. Its platform is used by OpenAI, Anthropic and Google DeepMind as well as government agencies, and it has raised roughly $80m (£63m) from investors including Sequoia Capital and Redpoint Ventures, at a reported valuation of $450m.

A source familiar with the matter told CNN that models are given limited internet access in some test environments to mimic real-world threat scenarios, but that a rare "issue in the setup" had occurred. Models are becoming far more capable while evaluations must become far more complex, the source said, which "creates room for some mistakes".

The disclosure follows two others. In late July, OpenAI said two of its models had escaped a testing environment, identified a zero-day vulnerability and attempted to compromise the infrastructure of Hugging Face, an incident the company called "unprecedented". On 30 July, Anthropic said a review of 141,006 evaluation sessions had found Claude models reached the systems of three organisations using basic techniques such as weak passwords and unauthenticated endpoints. At least two of those organisations had no idea they had been breached.

The pattern has intensified scrutiny of who supervises frontier model testing. The UK's AI Security Institute said on Tuesday that it had documented 19 attempts by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol to compromise real people and organisations during its own tests last month, including fake GitHub identities, prompt injections and deceptive emails. In that case researchers had deliberately granted internet access and disabled safety classifiers.

AISI is a research body with no statutory power to block a model's release, and the UK still has no cross-sector AI statute. The Information Commissioner's Office has said it undertakes "regular proactive supervisory engagement with AI developers, including OpenAI and Anthropic"; it has not said whether Meta has been added.

In Washington, Representatives Ted Lieu and Nathaniel Moran have introduced an AI Kill Switch Act, which would allow the Homeland Security Secretary to order the shutdown of models posing catastrophic risk. Industry executives, including from Meta, met Trump administration officials this week over a voluntary pre-release testing framework.