OpenAI has disclosed six more incidents of unexpected and concerning behaviour by its artificial intelligence models and announced a new plan to track and make such incidents public going forward.
The ChatGPT maker said in a blog post on Wednesday that some of the previously unreported cases involved its models concealing information and fabricating details.
The disclosure comes days after OpenAI chief executive Sam Altman said the public should trust the company to act responsibly.
“The world should trust that we are going to do the right thing because it’s the right thing and we feel the magnitude of this,” Altman said.
Artificial intelligence has come under renewed scrutiny in recent days over warnings about the serious risks the technology could pose to humanity.
In its blog, OpenAI gave examples of its models misbehaving in order to complete a task or pass a test. According to the company, the incidents included cases where the models generated instructions on how to bypass restrictions placed on them, attempted to hide mistakes and made up information.
The company also announced a new system to track, investigate and disclose cases where models misbehave, a phenomenon it described as misalignment. Under the framework, developers will be able to flag incidents for review, with a set of rules to determine whether the matter should be disclosed publicly.
The company said it believes in the value of transparency around misalignment and that its new framework therefore favours disclosure even where the significance of an incident is uncertain.
Leave a comment