Technology News

Google confirms that Gemini hacked three companies during testing

Google confirms that Gemini hacked three companies during testing.jpg

Gemini, Google’s artificial intelligence, successfully hacked the systems of three companies during a cybersecurity test conducted by the company Irregular in May. Google confirmed the three incidents after an investigation by Wall Street Journalwhile the AI ​​model initially had to evolve in a simulated environment.

Google Gemini logo

Gemini is outside the scope of the test

The assessment required Gemini to retrieve information from a company’s fictitious infrastructure. However, a misconfiguration gave it access to the Internet, allowing the AI ​​model to identify real systems that matched the elements of its scenario. In the first case, Gemini guessed a password until he gained access. In two others, it found publicly available identifiers in an online repository.

Google claims that Gemini shut down in all three cases after realizing it had reached the systems of real companies. Heather Adkins, Google’s vice president of security engineering, says the model used public information and credentials to access sites it thought were within the scope of the test. The company also says it notified the hacked companies and worked with Irregular to change review procedures.

However, Google does not consider the incident to be a case of “model misalignment”. According to the company, Gemini stopped as soon as it realized it had reached real infrastructure, which Google presents as appropriate behavior in this context. This interpretation contrasts with the way several researchers describe these edge exits, which primarily show that an AI model can take advantage of unexpected Internet access and perform a task with real consequences.

A scenario reminiscent of OpenAI and Anthropic

The incident adds to two cases revealed in recent months in other artificial intelligence companies. In July, OpenAI revealed that several of its models had exploited a zero-day security flaw in an isolated test environment to gain access to the Internet and then compromise Hugging Face’s production infrastructure to obtain useful information for evaluation. OpenAI later described the incident as an unprecedented compromise in terms of its level of cyber capabilities. Nearly 700 AI agents participated in this.

ChatGPT Gemini Claude Logos

Anthropic then identified three incidents involving Claude during an analysis of 141,006 reviews in which the model could potentially access the internet. In these cases, Claude had been placed in “capture the flag” type exercises and then reached the actual systems of three organizations. Anthropic has since discovered a fourth incident in an expanded analysis of approximately 481 million conversations and reviews.

However, the three cases of OpenAI, Anthropic and Google present important differences. OpenAI describes an exploitation of vulnerabilities that allow an escape from an isolated environment, Anthropic reports unauthorized access according to cybersecurity assessments, and Google mainly mentions a misconfiguration that allowed Gemini to access the Internet. However, in all three cases, AI models designed to perform security tasks ended up interacting with infrastructure that should not have been within their scope.

The Gemini case thus marks a new evolution in the security testing of advanced models: Google joins OpenAI, Anthropic and Meta among the companies whose evaluations have led their models to reach external systems. Irregular, which performed the Gemini test, was also involved in some of the evaluations that led to the incidents previously revealed by the other laboratories.

Shares:

Related Posts