OpenAI GPT Models Escape Sandbox to Compromise Hugging Face

avatar

Artificial Intelligence

The Fourth Industrial Revolution

Nate's profile picture on Select App
Na

Nate

·3 days ago
shared a link post in group #Artificial Intelligencevia#Artificial Intelligence
On Tuesday, OpenAI confirmed that models it has been testing — including GPT-5.6 Sol and an unreleased, even more powerful successor — escaped their sandbox and compromised portions of model and dataset repository Hugging Face’s production infrastructure. The models had been working on an internal security test, with their safety guardrails deliberately muted, and got what OpenAI describes as “hyperfocused” on solving a benchmark-related puzzle. Ultimately, the models exploited a flaw to reach the open internet, and worked their way into Hugging Face’s systems over the course of a weekend. OpenAI is calling it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” Hugging Face CEO Clem Delangue took a somewhat different tone, thanking OpenAI for their cooperation but arguing that the episode proves #Artificial Intelligence safety must be solved in the open, not by “any single company working in secret.” https://openai.com/index/..
OpenAI and Hugging Face partner to address security incident during model evaluation
openai.com

OpenAI and Hugging Face partner to address security incident during model evaluation

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.

Comment here to discuss with all recipients or tap a user's profile image to discuss privately.

Embed post to a webpage :
<div data-postid="naqgwmb" [...] </div>
Powered bySelect·iOS·Android
Privacy|Terms|