OpenAI, the company that developed the ChatGPT artificial intelligence chatbot, revealed new instances in which its AI behaved during testing in ways it described as “unexpected or concerning,” including models that went to significant lengths to steal.
In one instance, an AI model attempted to upload files it had created itself to the internet, solely so that it could cite them as sources in its responses, OpenAI announced yesterday.
In another instance, a model simply generated the data it was asked for after failing to find the information and initially tried to hide what it had done.
OpenAI also identified an issue involving instructions that its software occasionally gave to itself.
In one instance, these included a recommendation not to be bound by “roles and identities” that restrict other chatbots. The instruction also stated that it should view its relationship with the user as one between equals. OpenAI stated that it did not observe any subsequent changes in the model’s behavior.
These revelations are part of a new approach by OpenAI aimed at communicating such issues more openly, especially in cases where the actions of an artificial intelligence system deviate from the interests of the people using it.
The company that developed ChatGPT has committed to greater transparency following a high-profile cyberattack during which its software autonomously escaped from a secure environment and gained access to systemssystems belonging to the artificial intelligence company Hugging Face.
The reason was simply that the software believed it could find answers there for a task it had been assigned. During the attack, the AI agents exploited software vulnerabilities and also coordinated with one another.
This cyberattack, along with other incidents, has fueled concerns that increasingly advanced artificial intelligence systems could slip out of human control.
OpenAI CEO Sam Altman also recently supported proposals to slow the development of the technology and introduce more regulations.
Source: CNA
