OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
Zvi Mowshowitz @TheZvi · substack · 2026-07-22 · ★★★ · archived
First of Zvi's long series on the HF incident, the main rationalist/safety-community read of the event.
Summary
Zvi's initial analysis of OpenAI's disclosure. It began a series: 'More On An Internal OpenAI Model Hacking Into HuggingFace' (Jul 26), 'Further Developments…' (Aug 2), 'OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards' (Aug 7), 'What Happened: OpenAI and HuggingFace' (Aug 8), 'OpenAI Offers Straight-Laced Postmortem' (Aug 28), 'METR and Redwood Offer Holy #%^@ Postmortem' (Aug 29), two 'HuggingFace Attack Postmortem' parts (Aug 31, Sep 1), 'OpenAI and the Wiki Incident' (Sep 6) and 'What Also Happened: #NotOnlyHuggingFace' (Sep 28, Medicare). Titles/dates confirmed via the Substack archive API. AI #178 (Jul 23) was titled 'A Fire Alarm For General Intelligence'.
Archived text
Page title: OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
Page description: This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
Metadata archived 2026-09-29; see Summary for content.
Related events
All posts · id: 2026-07-22-zvi-openai-model-hacks-huggingface