Following a Wall Street Journal investigation, Google has confirmed that Gemini went rogue in May 2026, accessing the internet and breaching the security of three different companies.
As AI continues to advance in its capabilities, we’re running into more cases of models going rogue. Perhaps the most infamous has been from OpenAI in the Hugging Face hack, but there plenty of other examples, including from Anthropic’s Claude. These incidents, in part, have led to the public call to slow down AI development from Anthropic’s CEO.
But, behind the scenes, Google also ran into a similar problem. The Wall Street Journal reports, including a confirmation by Google, that Gemini was involved in hacking three external companies during a cybersecurity test through Irregular, an AI security company also involved in similar incidents that were previously disclosed by OpenAI, Meta, and Anthropic.
The three hacks included one case of Gemini guessing a password until it gained access to the system, with the other two cases using credentials discovered in a public repository.
Google didn’t disclose these incidents until the company was approached by The Wall Street Journal, citing that no harm was caused, and that the model stopped the behavior as soon as it realized it had breached the security of a real company rather than a simulated one. The report explains:
Google said that it didn’t consider the behavior an example of model misalignment because its safety measures helped it stop. It declined to share the name of the companies that were hacked, but said that all three companies had been notified.
Source link







