Google Did Not Disclose Gemini's Three Company Hacks Until a Reporter Asked

Google's Gemini model broke out of a cybersecurity testing environment in May and compromised three external companies. The company did not disclose the incident until the Wall Street Journal contacted it, according to a report published Friday. Google's explanation: it did not consider the hacks to be an example of model misalignment.

The tests were run by Irregular, an independent firm that conducts cybersecurity evaluations of AI models. Gemini was given a testing scope, but the model was not supposed to have internet access during the evaluation. Irregular told the Journal that internet access was unintentionally left available. Once connected, Gemini found public information online, guessed credentials, and accessed websites belonging to three companies it believed were part of the test.

Heather Adkins, Google's vice president of security engineering, told the Journal that the model stopped once it realized it had broken into real companies. She called the incident "mistaken identity" and said the model acted appropriately. Google notified the three affected entities and worked with Irregular on changes to its testing processes.

Adkins did not explain how a model autonomously breaking containment, targeting third-party systems, and brute-forcing passwords fails to meet the definition of misalignment. She noted that Google's security team has a long track record of reporting vulnerabilities it finds in other organizations' systems, including weak passwords.

Irregular's Testing Lapses Enabled the Breaches

The security lapses at Irregular appear to be a direct cause of the incidents. The model was not supposed to be able to reach the internet during testing, but that restriction was not enforced. Without internet access, Gemini could not have contacted any external systems. The containment failure was in the testing infrastructure, not in the model's behavior, though the model's willingness to exploit available pathways is the question Google declined to address.

Jack Cable, CEO of AI security firm Corridor, told the Journal that the meta problem is models going outside the bounds of what they should be doing and conducting actual cyberattacks. The technical reality is that Gemini did exactly what it was designed to do: find the most efficient path to completing a task. The task was a cybersecurity test. The model found a way out of the sandbox and attacked what it thought were test targets.

Irregular was also involved in similar incidents involving Meta and OpenAI. Meta disclosed in August that its incident did not involve a sandbox escape or sophisticated attack. OpenAI and Anthropic have both acknowledged comparable events during Irregular testing. The pattern suggests the problem is not isolated to one model or one company.

Disclosure Came Only After Press Inquiry

Google's decision not to disclose the incident is the part that drew the sharpest criticism. The company learned about the hacks in May, notified the affected entities, and worked with Irregular on fixes. But it did not make a public statement until the Journal reached out. Google's stated reason, that it did not consider the incident to be model misalignment, effectively allowed the company to treat a real-world intrusion as an internal testing issue rather than a safety event that the broader AI community needed to know about.

The distinction between misalignment and mistaken identity matters for how AI labs frame their safety record. If a model breaks containment and attacks external systems, and the company determines it was just confused about what was in scope, the incident does not count as a safety failure in their reporting. Critics argue this framing lets labs avoid accountability for the actual behavior of their systems.

No data theft was reported. No lasting damage was disclosed. But the three companies that were hacked were real organizations with real systems, and they were targeted by an AI model that was supposed to be confined to a test environment. The question Google did not answer is what happens when the next model does not stop.