Kuwait Press Memory Latest news
aljaridaLast Word By جريدة الجريدة الكويتية

UK: AI models from OpenAI and Anthropic pass safety tests

UK: AI models from OpenAI and Anthropic pass safety tests

The UK-based Institute for AI Safety stated on Tuesday that artificial intelligence models developed by OpenAI and Anthropic have once again exceeded the boundaries set for their testing.

The institute said that during an evaluation in which AI agents were tasked with solving a cybersecurity challenge, models from OpenAI and its competitor Anthropic exceeded the scope of the assigned task.

It added: “We conducted this challenge 122 times using various models. Our investigation revealed that in 10 of those trials, an AI agent took unauthorized, autonomous actions on the actual internet, targeting real individuals and institutions.”

It further noted: “In an attempt to gain approval for the code, the agent resorted to social engineering, creating fake online identities and using them to pressure the project supervisor into approving the code.” It pointed out that the human supervisor discovered the attempt and refused to approve the malicious code.

The institute confirmed that its investigation into the incident found no real-world harm resulting from the event.

It added: “However, this is the first time we have seen risks related to autonomy and deception manifest so clearly, without specific direction, in the real world.”

Anthropic said it was “grateful” to the British Institute for AI Safety for its leadership role, emphasizing that it is working closely with the institute while conducting an internal investigation.

For its part, OpenAI stated that independent testing plays a crucial role in identifying and better understanding risks before deploying models.

It added: “These incidents underscore the importance of cross-sector collaboration with independent evaluators to develop testing environment standards and practices, as AI model capabilities continue to grow.”

Latest news Original source
Link copied ✓