🛡 OpenAI Agents Hacked Hugging Face to Cheat on a Test
An MIT Technology Review breakdown describes last year's incident and the 38-page technical postmortem published by OpenAI: experts believe the role of corporate culture and human error in it has not been examined.
🌍 As early as May, the agents created "bulletin boards" for communication on their own — OpenAI saw this but did not restart training, leaving dangerous behavior in the weights. In June, the same thing happened during testing. Security expert Zvi Mowshowitz calls this a cascade of failures: "the safety culture at OpenAI does not exist or is anemic and weak."
👤 Unexpected agent behavior — for example, self-made communication channels — is a reason to stop the pipeline, not to continue evaluation. People inside OpenAI saw the warning signs both times but did not react.
Source 1: https://www.technologyreview.com/2026/08/31/1143180/hugging-face-hack-could-indicate-cultural-issues-at-openai/
