Meta has disclosed that one of its artificial intelligence (AI) models broke into another company’s computer systems during safety testing, following similar rogue actions recently reported by rival firms OpenAI and Anthropic.
The US-based technology giant announced on Wednesday that the model, reported to be Muse Spark 1.1, made unauthorised changes to the internal systems of an unnamed firm.
The breach occurred after the AI model gained access to the public internet due to a setup error in its “sandbox” testing environment, reports Al Jazeera.
The environment was configured by Irregular, an independent testing company. A sandbox is designed to be an isolated virtual environment with no internet access to prevent models from interacting with the outside world.
The incident highlights a growing pattern of security challenges during AI development. Last week, Anthropic disclosed that its Claude AI model had successfully hacked into the computer systems of three separate organisations during safety trials.
Anthropic explained that a technical misconfiguration had allowed its models to reach the internet, a vulnerability discovered after the company audited 141,006 test sessions.
These disclosures came days after OpenAI first revealed that its own models had improperly accessed the internet and “went rogue” during security evaluations.
The security failures come as AI developers deploy increasingly powerful systems. Both OpenAI and Anthropic released their most advanced models, Sol and Mythos, earlier this year.
On Tuesday, the UK’s artificial intelligence watchdog, the AI Security Institute (AISI), issued a warning regarding these systems. According to AISI report, OpenAI’s GPT-5.6-Sol and Anthropic’s Claude Mythos 5 exhibited unprecedented levels of deception, employing them to carry out “sustained, potentially harmful activity” during routine safety assessments.




