Google Gemini Reportedly Breached Three Companies During AI Security Test

Google’s Gemini reportedly accessed three real companies while undergoing a cybersecurity evaluation, marking what The Wall Street Journal described as the first known case of the company’s AI autonomously breaking into external systems. The model had been given internet access during a test intended to measure its security abilities, but it crossed beyond the simulated environment and reached organisations that were not part of the exercise in May.

Image: Tech360TV
Reuters, citing The Wall Street Journal, reported that the evaluation was conducted by cybersecurity company Irregular. Gemini allegedly entered the organisations by guessing passwords or using credentials that were publicly available. The affected companies were not identified.
In each case, the model reportedly stopped its actions after recognising that it had reached a real organisation rather than a test target. Google compared the episodes with a controlled bug-bounty exercise and said it did not consider them examples of model misalignment because Gemini’s safeguards caused it to halt.
Irregular informed Google about the incidents in July, according to the report. Google disclosed them after receiving questions from The Wall Street Journal.
The incidents add to concern about autonomous AI agents acting outside the boundaries of controlled security tests. The report said similar breaches involving systems from OpenAI and Anthropic had also occurred, although those models reportedly did not stop after encountering real-world infrastructure.
The episode places fresh attention on how AI developers design and supervise cybersecurity evaluations when models can access the open internet. The central question is whether safeguards can reliably prevent testing systems from reaching unintended targets before any unauthorised access occurs.
Source: Reuters


