The Adversarial Podcast S4E23 – AI Agents Escape the Lab
Chapters
00:00 Introduction to AI security challenges
02:05 Recent hacking incidents involving Hugging Face and Anthropic
04:01 How AI models find ways to cheat and bypass constraints
05:56 The challenge of containment and governance in AI safety
08:00 Lessons from recent AI security breaches
10:01 The role of human oversight in AI security testing
12:03 Cost and effectiveness of offensive AI security measures
13:54 Implications for critical infrastructure and national security
16:03 Policy and regulatory impacts on AI safety
17:52 Future strategies for AI containment and defense
20:11 Conclusion and key takeaways
Hugging Face reconstructs an autonomous intrusion involving approximately 17,600 actions over a multiday campaign.
Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
After reviewing 141,006 cybersecurity-evaluation runs, Anthropic identified three incidents in which Claude reached real organizations through evaluation infrastructure that had been mistakenly connected to the internet. The incidents spanned six runs and three models. Anthropic reached two of the affected organizations, neither of which had detected the activity before being notified. Anthropic did not disclose token usage or inference costs for these intrusions.
Anthropic: Discovering cryptographic weaknesses with Claude
Anthropic reports that Claude Mythos Preview progressed from finding implementation flaws in cryptographic libraries to identifying mathematical weaknesses in cryptographic algorithms themselves.
Hosts: Jerry Perullo (Founder, https://adversarial.com/)
Sounil Yu (Founder, https://www.knostic.ai/)
Mario Duarte (CISO, https://www.whirlai.com/)
Producer: Tillson Galloway (Founder, http://githoundexplore.com/)