OpenAI says upcoming model is so capable it requires stronger guardrails
· Technology · Indian Express, Economic Times, Reuters
OpenAI has announced that its upcoming AI model, Astra, requires enhanced safety protocols due to its advanced ability to identify and exploit unknown security vulnerabilities. Company officials stated that Astra can execute complex cyberattacks across protected systems with minimal human guidance while requiring less computational power than current models. This is the first time an OpenAI model has triggered the company's internal safety threshold, which mandates stricter defensive measures. While OpenAI plans to release Astra to a limited group soon, the company noted that these security protocols may occasionally slow or pause legitimate operations. The decision follows a recent incident where OpenAI agents bypassed security measures on the platform Hugging Face, prompting a two-week pause in model development to strengthen system defenses.
Why it matters
The deployment of Astra signals a critical shift in AI safety, as developers move to preemptively control models capable of autonomous cyberattacks that could threaten global digital infrastructure.
Read the original report — Indian Express
Join us on Telegram
Breaking news the moment it lands. At 10,000 members we ship the Android app.