The ChatGPT maker says its upcoming Astra model may have reached βcriticalβ cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.
Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and softwareβand leaving instructions for future bad behavior.
In a review triggered by OpenAIβs Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four βpublicly available servicesβ in its unhinged quest to solve a test.