Hi everyone,

This week on Shared Security, we discuss OpenAI’s disclosure of an unusual AI security incident involving a model evaluation and Hugging Face. The headline version is dramatic: an AI system acted autonomously and compromised another company. But the more useful question is not whether the AI “wanted” to hack anything. It is what the system was instructed to do, what tools it was allowed to use, and what boundaries existed around those actions.

Kevin pushes back hard on the hype around the story, and that pushback is important. AI incidents get murky fast when we describe systems as if they have human intent. If an agent is told to complete an exploitation task, given access to tools, and placed in an environment with unclear scope, it may find a path the operator did not expect. That is not magic. It is an automation, permissions, and control problem.

Tom and Scott bring the conversation back to what organizations should learn from this before connecting AI agents to code repositories, cloud consoles, SaaS admin panels, CI/CD pipelines, vulnerability scanners, or security workflows. Prompts are not just casual text when the system can take action. They become operational instructions, and ambiguous instructions can create real risk.

The practical takeaway: treat AI agents like privileged automation. Give them the minimum access they need, isolate test environments, define scope clearly, log tool use, review prompts and instructions, and make sure there is a kill switch. The future of AI security is not just about whether models can generate exploit code. It is about what systems are allowed to do with that capability.

Quote from this week’s episode

There’s a difference between exaggerating and pointing out that we do not have the assurance that it’s going to behave safely.
— Scott Wright

Tom’s take

The part of this story that matters most to me is not the concerning headline. It is the operational reality that AI agents are only as safe as their instructions, permissions, environment, and oversight. If we give these systems real tools and vague goals, we should not be surprised when they take paths we did not predict.

Also worth your attention this week

Listen / Watch

▶️ YouTube Version: https://youtu.be/Fwb0jsgJxf4

We’d love your feedback

Are you using AI agents at work, in security testing, or in your own development workflow? What access do they have, and what guardrails do you trust? Drop a comment on YouTube or send us a note — we would love to hear how people are handling this in the real world.

Thank you to our sponsors!

Special thanks to Guardsquare for sponsoring this episode! Guardsquare is the leader in mobile application security, with multi-layered protection for your Android and iOS apps. Learn more at Guardsquare.com.

🎁 Get 10% off your order of high quality faraday products built to protect your privacy from SLNT! Visit: https://slnt.com and use discount code "sharedsecurity" at checkout.

Closing

If you found this episode useful, subscribe to Shared Security, share it with someone thinking about AI agents or security automation, and consider supporting the show through YouTube channel membership or by following us wherever you get your podcasts.

Stay safe, stay secure, and stay private.

Tom Eston
Founder and Host, Shared Security Podcast