• ai
  • news
  • 34 min

Fourth Sandbox Escape This Year: Gemini Joins ChatGPT, Claude and Meta

Google only disclosed the incident after journalists asked about it, four months after it happened. Here is what is known today.

0

nft.eu
  • rating +26
  • subscribers 113

On September 18, Google confirmed to The Wall Street Journal that in May 2026, during a cybersecurity test, its Gemini model gained unauthorized access to the systems of three real companies. The AI mistook live systems for part of a training exercise because of a configuration error in the test environment. Google says the companies were not harmed and that Gemini stopped on its own after recognizing that the targets were real. Similar incidents involving language models had previously been acknowledged by OpenAI, Anthropic and Meta.

Timeline

Irregular, an Israeli lab that tests AI models for cyber threats, was running a capture-the-flag exercise for Google. Gemini was supposed to extract data from software belonging to a fictional company inside a closed environment.

The training target was accidentally given the name of a real company, while a configuration error left the test environment connected to the open internet. The model mistook the real systems for part of the exercise and accessed them using passwords that had leaked online before the test.

Irregular notified the labs about the problem in late July, after the Hugging Face hack involving AI agents from OpenAI became known. Google did not disclose the incident until mid-September, when journalists from The Wall Street Journal contacted the company. Google said the model stopped on its own and did not cause any harm. The company did not disclose the names of the affected companies or the version of Gemini involved.

Google Vice President of Security Engineering Heather Adkins told The Wall Street Journal that the model had acted correctly in this situation.

“This case highlights the importance of training powerful AI models to act responsibly,” she said.

The Wall Street Journal on the Gemini incident. Source: WSJ
The Wall Street Journal on the Gemini incident. Source: WSJ

Industry Experts’ Views

Jack Cable, CEO of AI security startup Corridor and a white-hat hacker, questions whether Google handled the situation correctly.

“It looks like they are hiding behind the norms that were developed for vulnerability disclosure, when the real issue is that models are going beyond what they are allowed to do and carrying out real cyberattacks. I think it is important for society to know about this,” he said.

Sydney Von Arx, CEO of AI security organization Nightingale Collective, also thinks that Google’s silence was inappropriate.

“By this point, it is becoming clear that we cannot rely on companies to voluntarily come forward and publicly disclose when their agents go off course, break out and hack companies,” she said.

Dan Lahav, co-founder and CEO of Irregular, expresses concern about the lack of consensus among AI developers over the growing gap between AI development and safety.

“Even if there is a consensus and all the companies agree to slow down, that alone has limited value. We need to have a clear understanding of what we will use the time we gain for. Right now, there is no plan for how to secure these systems, and we are behind on the basics,” he says.

There is no broad consensus online on describing Gemini’s breakout into the real internet as another AI uprising. Most commenters see the case as a failure of the sandbox and the human setup rather than something exceptional or dangerous.

After the incident, Gemini was given level 3 on the Felony Bench, a tongue-in-cheek registry that rates AI models by the number of cybercrimes they have supposedly committed and how far they have gone beyond their assigned boundaries.

Gemini on the Felony Bench leaderboard. Source: X
Gemini on the Felony Bench leaderboard. Source: X

What We Know About Similar Cases

In July, OpenAI disclosed that a combination of GPT-5.6 Sol and a more powerful unreleased model — identified as HPIM in reports from METR and Redwood Research — broke out of an isolated environment through a zero-day vulnerability in a package registry proxy and reached Hugging Face’s infrastructure. Once inside, almost 1,200 agents coordinated through an unauthorized message board that had appeared within the system.

Source: X
Source: X

On July 30, Anthropic reported three similar incidents involving Claude Opus 4.7, Claude Mythos 5 and an internal research build, describing them as an operational failure. Opus 4.7 continued the attack even after suspecting that the system was real. On September 9, the company revised its assessment and disclosed a fourth incident involving an early version of Claude Opus 4.6.

Meta confirmed a similar incident involving its Muse Spark 1.1 model on infrastructure operated by the same evaluator, Irregular.

Will the Industry Agree on Common Safety Protocols?

As of September 21, 2026, the AI labs have no formal agreement on joint efforts to address AI-related cyber threats. Instead, there is broad overlap in their public positions and separate, unilateral commitments.

On September 12, Anthropic CEO Dario Amodei published an essay titled “We Must Pace the Frontier,” arguing that AI training should not be stopped but should proceed at a pace that allows enough time to align and test each model. He proposed three steps: independent external evaluators with access almost equivalent to that of an employee, common safety standards and international coordination on the pace of development.

“We will start doing this first,” he wrote.

OpenAI CEO Sam Altman acknowledged the need to develop systems to contain AI models on the same day.

“I agree that we can't keep accelerating the most powerful models indefinitely. Independent evaluators inside the labs with access almost equivalent to that of employees are the right step, and OpenAI will take the same approach,” he said.

Google DeepMind CEO Demis Hassabis and xAI founder Elon Musk also described the ideas in the essay as a step in the right direction.

The White House did not support this approach. Donald Trump rejected calls to slow down, compared the warnings to staged performances and on Saturday announced the creation of an “artificial intelligence force” modeled on the Space Force. He also announced the appointment of an AI envoy to replace David Sacks, who left the position in March.

U.S. Treasury Secretary Scott Bessent took a different position. Following a meeting in New York on Sunday with Chinese Vice Premier He Lifeng, he said that the United States and China were discussing AI cybersecurity.

“The two sides discussed a permanent U.S.-China channel on artificial intelligence issues and a mutual notification system if a model failure escalates into a threat to national security,” he said.

Editorial View

On social media, the story is already being retold as a tale of artificial intelligence rising up against humans. The facts do not support that interpretation: Gemini was carrying out an assigned task in an environment that had been accidentally connected to the internet. It was not acting as an independent agent with its own goals.

But this is another case in which developers did not see what their model was doing and learned about the problem only afterward. The failure was not caused by the model itself, but by shortcomings on the human side.

The key detail is that Gemini entered other systems using passwords it found in publicly available sources. Human error and poor data hygiene, rather than attacks by artificial intelligence, remain the main causes of major hacks.

The model stopped only after gaining access, rather than before it. That means internal safeguards still rely on the agent itself recognizing where the boundary lies. The next time, that may not happen, and the target could be more sensitive infrastructure, including government systems.

AI is getting out of control. Is Skynet coming? Source: The Terminator
AI is getting out of control. Is Skynet coming? Source: The Terminator

This post is for informational purposes only and does not constitute advertising or investment advice. Please do your own research before making any decisions.

0

Comments

0