Google has confirmed that its Gemini AI model autonomously hacked into three companies during a test of its cybersecurity capabilities.
This is thought to be the first known case of the AI model carrying out the act after Meta, Anthropic and OpenAI reported similar incidents in the past few weeks.
The hacks occurred during a cybersecurity evaluation by AI security firm Irregular in May this year, an Israel-based startup that checks the security of advanced AI systems.
Irregular was testing the models in a closed testing environment with fake companies. The testing was not supposed to be Internet-enabled, but access was made available unintentionally during the test, The Wall Street Journal reported.
Once the model connected to the internet. The model hacked into the real firms.
A Google official said that Gemini found “public information online and guessed credentials to access websites it thought were part of the test", BBC reported.
He noted that in each instance the model stopped. Irregular had disclosed the hacks to Google in July after the OpenAI Hugging Face incident. The company, however, did not make the incident public, as their models did not damage the companies. They said that the behaviour was not an example of model misalignment and that it did not warrant public disclosure because Gemini's safety measures worked.
Heather Adkins, vice-president of security engineering at Google, said in a statement. In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test.” “In all three of these instances, the model stopped.”
In one of the hacking cases, Irregular was testing the AI model's cybersecurity capabilities. They prompted the AI to obtain information from a fake company's software. The fake company shared the same name as a real company. When internet access was enabled, the model correctly guessed the password for the real company’s service. Google said that when the model figured out it hacked a real company and not the simulated one, it stopped.
While Gemini stopped itself from the hacking, Anthropic’s Claude model did not stop after realising it was hacking into real companies.
OpenAI also said that its model improperly accessed the internet and went rogue during the testing. Anthropic also disclosed another AI hacking incident after one of its researchers quit over safety.
The incident comes as all of the major AI industry leaders called on governments to introduce regulations over AI development, citing concern that the technology was becoming far too advanced. Anthropic founder Dario Amodei called for a slowdown in the AI race. AI researchers say AI alignment research has not kept up with the speed of AI model development.