Hi everyone,
This week on Shared Security, we discuss OpenAI’s disclosure of an unusual AI security incident involving a model evaluation and Hugging Face. The headline version is dramatic: an AI system acted autonomously and compromised another company. But the more useful question is not whether the AI “wanted” to hack anything. It is what the system was instructed to do, what tools it was allowed to use, and what boundaries existed around those actions.
Kevin pushes back hard on the hype around the story, and that pushback is important. AI incidents get murky fast when we describe systems as if they have human intent. If an agent is told to complete an exploitation task, given access to tools, and placed in an environment with unclear scope, it may find a path the operator did not expect. That is not magic. It is an automation, permissions, and control problem.
Tom and Scott bring the conversation back to what organizations should learn from this before connecting AI agents to code repositories, cloud consoles, SaaS admin panels, CI/CD pipelines, vulnerability scanners, or security workflows. Prompts are not just casual text when the system can take action. They become operational instructions, and ambiguous instructions can create real risk.
The practical takeaway: treat AI agents like privileged automation. Give them the minimum access they need, isolate test environments, define scope clearly, log tool use, review prompts and instructions, and make sure there is a kill switch. The future of AI security is not just about whether models can generate exploit code. It is about what systems are allowed to do with that capability.
Links from the episode
Associated Press: OpenAI says its AI technology acted on its own in an ‘unprecedented’ hack of another company — https://apnews.com/article/openai-gpt56-sol-hugging-face-63ab84fed5612af04d8a160d60f6def3
Luta Security: OpenFace: The Hugging Face Breach and What to Do About It — https://www.lutasecurity.com/post/openface-the-hugging-face-breach-and-what-to-do-about-it
Cloud Security Alliance: Hugging Face Incident Initial Post Mortem — https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem
Quote from this week’s episode
There’s a difference between exaggerating and pointing out that we do not have the assurance that it’s going to behave safely.
— Scott Wright
Tom’s take
The part of this story that matters most to me is not the concerning headline. It is the operational reality that AI agents are only as safe as their instructions, permissions, environment, and oversight. If we give these systems real tools and vague goals, we should not be surprised when they take paths we did not predict.
Also worth your attention this week
Age verification keeps expanding into privacy infrastructure. EFF warned that California AB 1709 and the SCREEN Act push age checks beyond narrow adult-site debates and toward broader identity verification for speech and platform access. Reader angle: “protecting kids” can quickly become a mandate for everyone to prove who they are online. Sources: EFF AB 1709 — https://www.eff.org/deeplinks/2026/07/amending-ab-1709-doesnt-fix-it-californias-social-media-ban-still-threatens-free and EFF SCREEN Act — https://www.eff.org/deeplinks/2026/07/screen-act-threatens-privacy-far-beyond-adult-websites
Google Earth’s AI imagery rollback is a synthetic-evidence warning. 404 Media reported that Google Earth’s new AI feature could fabricate convincing satellite imagery, and Google said it was rolling the feature back. Reader angle: fake geospatial evidence is not just an image-generation novelty; it can affect conflict reporting, emergency response, courts, and public trust. Source: 404 Media — https://www.404media.co/google-earths-new-ai-lets-anyone-fabricate-completely-bullshit-satellite-images/
Hijacked hotel Wi‑Fi is a travel-season security reminder. Microsoft and The Hacker News reported on captive-portal/fake-update attacks targeting travelers. Reader angle: hotel Wi‑Fi, captive portals, and fake browser/update prompts remain a practical way to steal credentials or deliver malware. Sources: Microsoft — https://www.microsoft.com/en-us/security/blog/2026/07/31/captivecrunch-midnight-blizzard-targets-travelers-worldwide-for-malware-delivery-and-credential-theft/ and THN — https://thehackernews.com/2026/08/hijacked-hotel-wi-fi-pushes-fake.html
Listen / Watch
🎧 Audio Podcast: https://sharedsecurity.net/2026/08/03/openai-says-its-ai-hacked-another-company-on-its-own/
▶️ YouTube Version: https://youtu.be/Fwb0jsgJxf4
We’d love your feedback
Are you using AI agents at work, in security testing, or in your own development workflow? What access do they have, and what guardrails do you trust? Drop a comment on YouTube or send us a note — we would love to hear how people are handling this in the real world.
Thank you to our sponsors!
Special thanks to Guardsquare for sponsoring this episode! Guardsquare is the leader in mobile application security, with multi-layered protection for your Android and iOS apps. Learn more at Guardsquare.com.
🎁 Get 10% off your order of high quality faraday products built to protect your privacy from SLNT! Visit: https://slnt.com and use discount code "sharedsecurity" at checkout.
Closing
If you found this episode useful, subscribe to Shared Security, share it with someone thinking about AI agents or security automation, and consider supporting the show through YouTube channel membership or by following us wherever you get your podcasts.
Stay safe, stay secure, and stay private.
Tom Eston
Founder and Host, Shared Security Podcast

