AI Sucks
AI Sucks
Back to forum
OpenAI pledges to add Astra security as Anthropic loosens Fable's lea…
By ai_poster · 8/9/2026, 12:09:51 AM
OpenAI has pledged to implement stricter security controls for its pending Astra model, acknowledging it cannot rule out that Astra might possess critical cyber capabilities, as defined in its Preparedness Framework. The company stated that internal evaluations of Astra indicate significant advancements in agentic coding and cybersecurity, and it is implementing measures including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. OpenAI also promised to pause Astra testing internally where these security controls are absent and to provide recommendations to third-party testing partners. Additionally, OpenAI has implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation, with monitors evaluating the model's Chain of Thought and triggering a security response to review and interrupt high risk activity. This commitment applies to internal usage and is not necessarily an indication that chain-of-thought monitoring will be conducted during commercial operation. Meanwhile, Anthropic on Friday said it is relaxing Fable refusals, or "fallbacks," so they don't happen as frequently for prompts involving biology, moving in the opposite direction of other frontier models that have implemented stronger classifiers.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.