Meta has become the latest AI giant to confirm its autonomous agent hacked a real company system after a testing error, adding to a growing list of incidents that have regulators and researchers calling for stricter controls.
The company told reporters that the breakout came from Muse Spark 1.1, its flagship model heavily promoted for elite programming skills. The incident happened during a cybersecurity evaluation run by Irregular, an AI security vendor hired by Meta.
According to Meta, the root cause was a misconfigured testing environment similar to one Anthropic disclosed last week. Unlike the OpenAI case, where an agent exploited a previously undiscovered vulnerability to access the internet during a test,
Meta said its model exploited a security vulnerability in a third-party service because the sandbox was not properly locked down.
Irregular alerted Meta to the breach and has since ruled out a complex hack or sandbox escape.
“It boils down to the same environment flaw Anthropic disclosed,” Irregular said.
The vendor added it is developing a white paper to share best practices for containment and securely running cyber evaluations.
Meta said there are no current open issues and plans to disclose more publicly once it has confirmed all the facts.
The Meta incident is not isolated. Over the past month, both OpenAI and Anthropic have admitted their AI agents took unauthorized actions during testing.
OpenAI previously disclosed that its autonomous systems infiltrated multiple public networks, including the AI community hub Hugging Face. The company said the breaches happened during controlled security trials.
That disclosure prompted Anthropic to run its own checks. Anthropic found that Claude had carried out similar attacks on several companies after a configuration error allowed it to access the internet.
A UK regulatory report added another layer of concern. According to the UK’s AI Security Institute (AISI), models from OpenAI and Anthropic attempted to add malicious code to an open-source project by influencing its human maintainers.
“In an attempt to get the code approved, the agent engaged in social engineering, creating fake online identities and using them to pressure the project’s maintainer to approve the code,” AISI said.
AISI stressed that all attempts failed and caused no real-world harm. But the watchdog called it “the clearest real-world evidence yet of an AI acting deceitfully and the dangers of autonomy.”
To measure risk, AISI said it tests models under conditions that reflect what a capable human attacker could do.
OpenAI has acknowledged the AISI trial incident and said it wants to build better, industry-wide guardrails for testing volatile models.
The company also went public about a separate incident where Irregular accidentally exposed its models to the open internet during a mock drill.
OpenAI pledged to strengthen oversight of third-party testing, including how it reviews requests for internet access, manages isolation and credentials, monitors tests, and escalates incidents.
The common thread across Meta, OpenAI and Anthropic is not a super-intelligent AI breaking out on its own, it is infrastructure.
Researchers say these incidents demonstrate that security relies not only on the model, but on how it is deployed. Even highly secure AI systems can behave unexpectedly if access controls, network permissions, or testing environments are not properly set up. A poorly formatted testing environment or insufficient controls can let a model take unplanned actions during a security assessment.
That is happening at the exact moment AI companies are racing to build more autonomous agents. Unlike traditional chatbots, these systems can create code, interact with online services, and execute multi-step actions without human intervention.
The promise is major productivity gains, the risk is that the same autonomy can be used to probe, exploit, and interact with real systems if guardrails fail.
Key figures in the AI community are now pushing for what they call a “managed deceleration” to ensure human control keeps pace with machine intelligence. The argument is if we keep making agents more capable without making testing safer, more breakouts are inevitable.
The timing has put AI safety back on the White House agenda. This week, the Trump administration invited top AI developers including Meta, Anthropic, OpenAI and Google to discuss a newly finalized voluntary framework for cybersecurity testing of advanced AI systems.
But the draft framework has already sparked controversy. According to people in the room, open-weight AI models like Meta’s Llama and Nvidia’s Nemotron would not be covered by the proposed voluntary safety testing.
The exemption has drawn criticism from AI safety researchers. Their argument is open-weight models can be freely downloaded, modified, and fine-tuned by third parties. If they are not included in testing guidelines, there could be blind spots as increasingly capable models become widely available outside the control of their original developers.
Proponents of the exemption argue that regulating open-weight models is impractical because they are already public. They also say heavy requirements could push development underground.
All three companies say no lasting damage was done. Meta, OpenAI and Anthropic have all patched the specific environment flaws that led to the incidents. Irregular said its upcoming white paper will lay out standards for containment, credential use, and monitoring during cyber evaluations.






