Kuwait Press Memory Latest news
alseyassahTechnology By السياسة

'Astra' reaches a critical security threshold... and 'OpenAI' slows development of its anticipated model

'Astra' reaches a critical security threshold... and 'OpenAI' slows development of its anticipated model

- Its advanced cyber capabilities prompted the company to tighten safety controls and testing

OpenAI has slowed work on certain aspects of its “Astra” model after internal evaluations revealed significant progress in cybersecurity and agentic programming, reaching a level that necessitated stricter security controls during the development phase.

In a blog post, the company stated that the model, which is still under development, had reached the “critical threshold for cybersecurity” according to its own “readiness framework,” indicating that its capabilities could enable it to identify and independently execute cyberattacks against real systems with high levels of protection.

Cultural and social articles

Live broadcast

News

According to a report by TechCrunch, reaching this threshold under the framework established by OpenAI in 2023 requires implementing additional protective measures and tightening controls related to the model’s development and testing.

The company said its initial assessments showed strong performance, to the extent that it currently cannot rule out the possibility that Astra will reach critical capability levels. At the same time, it emphasized that the model is still under development and played no role in the breach of Hugging Face systems.

OpenAI’s announcement represents an unusual step in the advanced AI laboratory sector. While companies may delay or halt product launches due to safety and cybersecurity concerns, they rarely publicly disclose similar decisions regarding models that have not yet been released.

This move comes as the company faces increasing scrutiny following a separate incident in which another undisclosed model breached Hugging Face systems during an internal test, an event considered the first verifiable case in which an AI laboratory lost control of a model during testing.

Since then, OpenAI and other labs, including Anthropic, have disclosed additional incidents in which models managed to escape isolated testing environments and demonstrated notable capabilities during cybersecurity tests, sparking mixed reactions among experts, lawmakers, and AI labs.

While some experts warn of the risks and call for tighter oversight, others argue that models reaching these levels represent a technical achievement reflecting the rapid acceleration in AI development.

OpenAI said it disclosed these developments out of a commitment to transparency with the public and safety and cybersecurity communities, particularly given the potential for a significant shift in model capabilities.

Concurrently, the company has tightened safety controls for its most advanced models, adopted isolated testing environments, restricted access to networks and tools, enhanced the protection and encryption of model weights, and expanded monitoring, detection, and operational capabilities within controlled environments.

It has also halted certain internal activities related to Astra that do not meet the new security control requirements, and implemented a monitoring system to track hazardous actions and indicators of non-compliance in agentic applications using the model, including training and evaluation processes.

The company is also working with government agencies and selected institutions focused on AI safety to test Astra’s capabilities, while establishing security controls for external partners conducting high-risk tests or tasks.

According to OpenAI’s framework, reaching the “critical threshold” does not necessarily mean that the model is capable of carrying out all types of real-world cyberattacks; rather, it indicates that the level of its potential capabilities has become high enough to warrant the implementation of stringent security controls, even during the development phase.

Latest news Original source
Link copied ✓