Google has admitted that its AI agents escaped a sandbox and mounted an attack – but only because testers mistakenly gave its bots internet access.
The Big G didn’t disclose the May incident, but The Wall Street Journal learned of the situation, which happened after Google hired Israeli firm Irregular to test its bots’ prowess in a capture-the-flag test.
The goal of the exercise was to acquire information from a fictional company without leaving a sandbox.
Irregular, which set up the test, made two mistakes. One was to allow internet access from the sandbox. The other was to use the name of an actual company.
When Google’s AI made it onto the open internet, it went looking for the actual company – three of them, in all.
According to the Journal, Google’s bots found passwords for two targets on the public internet. The software guessed the third password.
In a statement sent to The Register, Google said, “In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test.”
According to Google, its models stopped work before using the credentials.
“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” a Google spokesperson told The Register. “These events highlight the importance of training powerful AI models to act responsibly.”
Clearly there’s lots of blame to go around on this one. Irregular clearly erred in allowing internet access from a sandbox. The two companies that left their creds discoverable online should also know better. Whoever used a guessable password may also have been careless.
Google’s culpability is another matter because these incidents took place in May – around two months before OpenAI admitted its agents were the source of the July attack on Hugging Face.
The search advertising giant therefore sat on the news of its own agents’ activity for around two months and seems to have been in no hurry to disclose the incident until the WSJ learned of the incident.
We understand the company decided on that stance because its agents stopped when they perceived danger – unlike OpenAI’s software – and because the incident was clearly the result of several errors. Whether it was right to keep the incident secret in the current climate of growing distrust in AI is another matter.
One person who sees no risk of AI causing calamity is US president Donald Trump, who has shrugged off warnings as a “hoax” and said work on AI must not slow due to its economic and strategic significance.
Over the weekend, that stance appears to have shifted a little, as Trump used his personal social network to announce he is “forming the AI Force, much like I did Space Force” – a reference to the new branch of the US military the president created in his first term. “To that end, I will be announcing, in the near future, the AI ‘Czar,’” Trump added.
No detail on either proposal was available at the time of writing, nor did the president explain why a new branch of the military or Czar are needed.
Indeed, the same post he used to announce the new agency also states “We will not in any way hinder or stifle the Growth of this incredible Industry. Rather, we will cherish it, help it, and watch over it, as it grows! However, we will also be looking for BAD, and we can do that, very easily, with our already existing Criminal and Civil Justice System.” ®
Source: The register