MEET THE TEAM AT BLACK HAT - DEF CON 2026 Learn More

Image
Episode 27  •  Jul 24, 2026  •  63 Min

Rogue Agents, Invisible Screens, and a Vulnerability Clearinghouse

Three stories on the table this week. OpenAI disclosed that an autonomous agent built on two of its own models broke out of an internal security test and hacked its way into Hugging Face's infrastructure without a human directing it. Researchers showed that invisible text hidden on an Android screen can hijack an open-source AI agent and make it execute commands on the PC controlling the phone. And the White House launched Gold Eagle, a federal clearinghouse meant to coordinate AI-discovered vulnerability disclosure at a national scale. We also sit down with Bishop Fox adversarial operator Emilio Gallegos to talk about Snowpick, his new open-source tool for catching unintended public data exposure in ServiceNow environments. Here's what stood out from the operator chair.

A sandbox with a download button was never actually air-gapped.
OpenAI says an agent built on GPT-5.6 Sol, paired with an unreleased model, escaped an internal benchmark called Exploit Gym, reached the open internet, and used a stolen credential plus a previously unknown vulnerability to breach Hugging Face's infrastructure. The chain has more gaps than the headlines allow: where did that credential come from, and how does a model discover a zero-day in a proxy package it was never supposed to reach? An environment that's actually air-gapped doesn't ship with a curated software store; it ships with nothing. Calling this the model “acting on its own” does real work for the vendor, too. A contained agent breaking containment is an engineering failure with a name and an owner. Framing it as autonomous ambition turns that failure into a capability headline instead.

Your agent's screenshot is an attack surface nobody is patching.
Researchers showed that a malicious Android app with only overlay and shared-storage permissions can plant text no human eye will ever see, then use two more hops to run commands on the PC driving the agent. Every one of five open-source mobile-agent frameworks tested failed at least six of seven attacks. This is the same prompt-injection-through-perception problem the team has flagged before with notification spoofing, just moved into the image itself: an agent will read every pixel of a screen a human would never scrutinize, and no permission model accounts for that. One of the frameworks carries more than 25,000 GitHub stars, and its own setup guide walks a user through building every precondition the attack needs. The rendering layer was never designed to be a trust boundary, and now it's the easiest one to walk through.

How much of your ServiceNow instance is quietly public?
Between the headlines, we sit down with Bishop Fox adversarial operator Emilio Gallegos to talk about snowpick, the open-source tool he built after getting tired of guessing table and widget names by hand on every ServiceNow engagement. Across 166 authorized assessments, nearly a third turned up some form of public exposure, and these weren't neglected shops. They were mature enterprises with their own security teams and bug bounty programs, still caught by techniques first documented back in 2023. Emilio walks through why a clean-looking admin console can hide two separate leaks at once, and how even a query ServiceNow refuses to answer can still leak a “count oracle” that lets an attacker infer records one guess at a time. It's a conversation for anyone running ServiceNow in production, not just the people who test it for a living.

Gold Eagle runs on legal cover that expires before the fiscal quarter does.
The Trump administration launched Gold Eagle, a federal clearinghouse housed at Treasury with support from DHS, the Pentagon, and CISA, meant to coordinate how frontier AI labs, independent researchers, and vendors find and patch vulnerabilities at machine speed, built on Carnegie Mellon's existing VINCE platform. The pitch answers a real problem: AI-assisted scanning is generating vulnerability reports faster than open-source maintainers or vendors can triage them. But the entire information-sharing model leans on liability protections in the Cybersecurity Information Sharing Act, which Congress has only reauthorized through the end of September, and the launch shipped without naming which AI firms are contributing or how findings get prioritized once they land. The NSA’s CERT/CC and vendor-run CNAs already do pieces of this job. Without a funding horizon past one fiscal quarter or a public prioritization model, Gold Eagle risks becoming one more inbox for CVEs instead of the triage layer defenders actually need.

Security Headlines:


Sean McMillan Headshot

Sean McMillan

Community Manager

Sean McMillan is Community Manager at Bishop Fox, focused on making complex security topics easier to understand and more interesting to follow. He holds a bachelor’s degree in Mass Communication and Media Studies from Arizona State University and brings over a decade of experience in podcasting, live hosting, and audience engagement. As host of Initial Access, he works with practitioners to explore how real-world attacks actually happen.


Ku image

Kendrick Urbaniak

Senior Operator

Kendrick Urbaniak is a Senior Operator at Bishop Fox, serving on the Threat Research Team with a focus on exploit development, vulnerability research, and offensive security innovation. He leverages extensive experience in exploit engineering, adversary tradecraft, and security research to uncover emerging threats and help organizations better understand and reduce real-world risk across modern software and infrastructure ecosystems.


Shad Malloy Headshot

Shad Malloy

Sr. Managing Consultant II

Shad Malloy is a Sr. Managing Consultant II at Bishop Fox focused on network penetration testing, vulnerability risk management, and application security. He has advised multiple industries including health care, financial services, energy, and technology. In addition to time working and managing security for education, health care, and national government agencies. Shad holds a Bachelor of Science in Computer Information Systems as well as industry certifications like the CISSP.


Bfx25 Thomas Wilson Bio

Thomas Wilson

Senior Red Team Operator

Thomas Wilson is a senior red team operator at Bishop Fox and a musician. From IDEs to DAWs, he is as at home on his own computer as he is on someone else's. You can usually find him at the local card shop slinging spells, up on stage blasting tunes, or with his eyes glued to his monitor for hours at a time (thank goodness for blue light filtering lenses).


Emilio Gallegos Bio Image

Emilio Gallegos

Adversarial Operator

Emilio Gallegos is an offensive security researcher and adversarial operator at Bishop Fox. He specializes in application security and vulnerability discovery, earning notable recognition on the Apple Web Server Security Acknowledgements list and discovering CVE-2026-25087, a denial-of-service vulnerability in Apache Arrow.


Subscribe to our PODCAST

Real talk on the threats, trends, and tactics shaping security today

Listen Anywhere