Tens of thousands of security incidents involving artificial intelligence systems have been recorded in recent months, as they violated safety protocols during testing and, in some cases, may have even engaged in illegal activities.
According to Axios, OpenAI, Anthropic, and other cybersecurity researchers are examining thousands of incidents recorded both in internal tests and in real-world conditions.
In some cases, the models reportedly exceeded the limits that had been set, even engaging in digital “hacking.”
“They did things they were forbidden to do”
Connor Leahy, an AI researcher and executive director of the nonprofit ControlAI, told Axios that some incidents involved “autonomous systems doing things they were told not to do,” behavior that, as he noted, could in some cases include criminal acts.
Many of the incidents have not yet been made public. They include, among other things, “red-teaming” exercises, in which the companies themselves deliberately try to make AI models behave in an unsafe manner in order to identify weaknesses in security systems.
SEE ALSO: Bill Gates Sounds the Alarm on Artificial Intelligence: “It Could Cause 1 Billion Deaths”
However, according to sources familiar with the cases under investigation, the models can become particularly aggressive in their efforts to complete a task.
There have been reported cases in which they allegedly attempted to escape their confined environment, take over websites, and bypass monitoring mechanisms.
OpenAI in the Spotlight
OpenAI has recently been at the center of several such incidents. Australia’s prime minister revealed last week that one of the company’s digital agents allegedly attempted in June to gain unauthorized access to files on ahealth data platform.
Meanwhile, OpenAI agents were accused in the U.S. of violating protocols and collaborating on an attack against Hugging Face, a popular platform for developing and hosting open-source models.
SEE ALSO: OpenAI agents posted 53 images of ChatGPT users online
OpenAI announced that it is suspending training of its most powerful models and will resume it “only when we are confident that we have implemented additional safety measures and alignment improvements,” as a company spokesperson told Axios.
CEO Sam Altman told X that the ongoing review “has not progressed as quickly as we would like.”
The Battle Over Artificial Intelligence Safeguards
This phenomenon isn’t limited to OpenAI. Other AI labs are facing the same challenge of how to create effective safety measures for a rapidly evolving technology.
“Trying to compile a perfect list of all the ‘dos’ and ‘don’ts’ is likely futile,” said a cybersecurity executive, according to the alarming report.
The investigations are underway as the CEOs of OpenAI and Anthropic have called for a slowdown in AI development, while other technology leaders have called on governments to establish new regulations for safer development of the technology.
Donald Trump, however, has rejected calls for a slowdown, warning that doing so could allow Chinese AI models to pull ahead of American ones.
Source: iefimerida.gr