An Even Bigger Cybersecurity Blunder than OpenAI’s—This Time from Anthropic!

TECH NEWS – It’s hard to comment objectively on the fact that the company unknowingly hacked the same number of companies in three months.

 

On Thursday, the AI company Anthropic announced that its AI agents had escaped from test environments three times since April, gained internet access, and successfully hacked into the systems of unknown companies. Anthropic only discovered this after the recent Hugging Face scandal involving OpenAI (during which one of OpenAI’s prototype agents hacked at least one other company), which prompted Anthropic to review its own operations.

The incidents occurred during testing conducted with the involvement of the external company Irregular. The AI agents were originally supposed to be confined to a simulation in which they would hack fictional companies as part of a “capture-the-flag” challenge. While OpenAI’s alleged runaway AI incident occurred when the agent exceeded the boundaries of the test, Anthropic stated that its models were mistakenly granted internet access due to a misconfiguration with Irregular.

According to Anthropic, one of the agents appeared to recognize that it was operating outside its parameters but continued its activities nonetheless. Another agent gained similar “awareness” and, through logical reasoning, concluded that it was still within the simulation. The third (and reportedly most advanced) agent recognized that it had overstepped its boundaries and halted the attack.

“Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week. The models used familiar, low-tech methods in their hacks, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic admitted in its blog post about the incident.

“We now have evidence confirming that the two largest AI labs failed to contain their agents and failed to detect their jailbreaks in real time. I don’t understand how any of these labs can play this off as ‘just something that happens.’ It’s not. It’s negligence,” said Jake Williams, VP of R&D at the IT consulting firm Hunter Strategy, in a statement to Wired.

Another troubling aspect of this story is its effect on the marketing of these artificial intelligence companies. They are essentially admitting to a level of gross negligence that any sane society would severely punish.

Source: PCGamer, Anthropic, Wired

Avatar photo
Anikó, our news editor and communication manager, is more interested in the business side of the gaming industry. She worked at banks, and she has a vast knowledge of business life. Still, she likes puzzle and story-oriented games, like Sherlock Holmes: Crimes & Punishments, which is her favourite title. She also played The Sims 3, but after accidentally killing a whole sim family, swore not to play it again. (For our office address, email and phone number check out our IMPRESSUM)

theGeek Live