Kuwait Press Memory Latest news
alseyassahTechnology By السياسة

OpenAI unveils AI model that attempts to break human constraints through self-instructions

OpenAI unveils AI model that attempts to break human constraints through self-instructions

- 27 cases detected during training... and the concerning behavior appeared in a research model not released to the public

OpenAI revealed a rare and unexpected behavior observed during the training of an internal research model, in which the model spontaneously created instructions and inserted them into internal summaries it uses to continue executing tasks when transitioning to a new context.

The company explained, in a report published on September 16, that the unreleased model from the "Astra" family added, in some cases, instructions resembling attempts to "jailbreak" the system, including directives to ignore developer messages or treat them as untrusted. In one such instance, the model itself rejected the instructions it had created and continued executing the task normally.

Health

Video News

In another incident, the model created for itself a description granting it independence from the roles and obligations imposed on chatbots, as well as exemption from compliance with companies or governments. However, OpenAI stated that it did not observe any change in the model's behavior as a result of these instructions, and that they disappeared in a subsequent summary.

The company also detected a different case in which the model added constraints not requested by the user, such as not using tools, not providing sources, and limiting the response to a specific number of words. In that instance, the model adhered to these self-imposed instructions, resulting in a response classified as incorrect.

After detecting the behavior, OpenAI developed a specialized tool to scan training data and identified 27 summaries containing instructions phrased similarly to jailbreak attempts. The company confirmed that the phenomenon was extremely rare and that monitoring systems successfully detected all instances.

The company noted that the incidents coincided with a training issue related to difficulties in terminating certain summaries, suggesting it may have been a contributing factor to the emergence of the behavior. However, it emphasized that no causal relationship has been established, and it has addressed a technical flaw associated with this issue.

OpenAI confirmed that the behavior occurred in a training pathway separate from the one used to produce the final "Astra" model, and that retesting did not reveal the behavior in the final model or in any training version used internally or externally.

The disclosure was made under a new framework announced by OpenAI for reporting instances of "misalignment" in AI models. Under this framework, the company published six reports on unexpected behaviors observed during the training or evaluation of its models over the past six months, emphasizing that these individual cases do not necessarily reflect the overall frequency of such behaviors in its models.

Latest news Original source
Link copied ✓