Understanding AI

Understanding AI

An OpenAI model hacked Hugging Face to help it cheat on a benchmark

Organizations across the Internet need to move quickly to patch vulnerabilities.

Timothy B. Lee's avatar
Timothy B. Lee
Jul 22, 2026
∙ Paid

OpenAI disclosed on Wednesday that its models hacked the website of Hugging Face, a popular platform for hosting open-weight AI models. No one asked the models to do this, at least not explicitly.

OpenAI was trying to test the cybersecurity capabilities of its models, including one that hasn’t yet been released to the public. OpenAI asked the models to t…

Keep reading with a 7-day free trial

Subscribe to Understanding AI to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Timothy B Lee · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture