The Adversarial Podcast S4E23 – AI Agents Escape the Lab

Chapters

00:00 Introduction to AI security challenges

02:05 Recent hacking incidents involving Hugging Face and Anthropic

04:01 How AI models find ways to cheat and bypass constraints

05:56 The challenge of containment and governance in AI safety

08:00 Lessons from recent AI security breaches

10:01 The role of human oversight in AI security testing

12:03 Cost and effectiveness of offensive AI security measures

13:54 Implications for critical infrastructure and national security

16:03 Policy and regulatory impacts on AI safety

17:52 Future strategies for AI containment and defense

20:11 Conclusion and key takeaways

HuggingFace: Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Hugging Face reconstructs an autonomous intrusion involving approximately 17,600 actions over a multiday campaign.

Anthropic: Investigating three real-world incidents in our cybersecurity evaluations

After reviewing 141,006 cybersecurity-evaluation runs, Anthropic identified three incidents in which Claude reached real organizations through evaluation infrastructure that had been mistakenly connected to the internet. The incidents spanned six runs and three models. Anthropic reached two of the affected organizations, neither of which had detected the activity before being notified. Anthropic did not disclose token usage or inference costs for these intrusions.

Anthropic: Discovering cryptographic weaknesses with Claude

Anthropic reports that Claude Mythos Preview progressed from finding implementation flaws in cryptographic libraries to identifying mathematical weaknesses in cryptographic algorithms themselves.

Hosts: Jerry Perullo (Founder, https://adversarial.com/)

Sounil Yu (Founder, https://www.knostic.ai/)

Mario Duarte (CISO, https://www.whirlai.com/)

Producer: Tillson Galloway (Founder, http://githoundexplore.com/)