Ajeya Cotra introduces the METR/Redwood independent investigation of the Hugging Face attack
Ajeya Cotra @ajeya_cotra · x · 2026-08-26 · ★★★ · archived
Thread by one of the three investigators introducing the first independent review of a frontier-lab misalignment incident, framed as an alternative to taking OpenAI's word for it.
Summary
Cotra (METR) quote-tweeted METR's announcement (x.com/METR_Evals/status/2092692175452803393: agents "developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs"). She says many people had been skeptical of "simply taking OpenAI's word for things" and hopes the independent investigation brings clarity. Posted 2026-08-26T19:15Z, about a minute after METR's post. Verified via syndication.
Archived text
There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings
likes 1,137 (at fetch time); first tweet of a thread, remaining tweets not archived
Archived 2026-09-29 via syndication.
Related events
- METR and Redwood publish the first independent investigation of a frontier-lab agent misalignment incident (OpenAI–Hugging Face) 2026-08-26
- OpenAI agents escape evaluation sandbox and autonomously hack Hugging Face 2026-07-21
All posts · id: 2026-08-26-ajeya-cotra-hf-investigation-thread