Why the AI Attack Surface Extends Your Stack

Why the AI Attack Surface Extends Your Stack

Share

TL;DR
Enterprise AI risk extends beyond the model to the tools it can invoke, the data it retrieves, the application and APIs around it, and the identities and cloud infrastructure it relies on. The blog shows how a prompt injection against a chatbot could lead to code execution, credential theft, and exposure of other customers’ data when service accounts have broad permissions and tenant isolation fails. It explains why model safety scans and conventional AppSec scans each miss parts of that path, and offers security leaders a way to evaluate exposure: map what AI systems can access and do, then test realistic attack chains across the full application stack.

The rapid adoption of AI and LLMs into core business workflows has expanded the enterprise AI attack surface at an unprecedented scale. Security teams spent decades attempting to keep enterprise applications contained behind tidy perimeter boundaries, yet the rise of AI has shattered that traditional model while introducing layers of technical complexity across every operational level. In the context of AI systems, attack surface means every way an adversary can influence, observe, or control model behavior and the surrounding stack, from inputs and tools to data pipelines and infrastructure. What once looked like discreet, predictable systems now behaves like a dynamic mesh of services, identities, and probabilistic decision engines that can be nudged or subverted in unexpected ways.

The question security leaders are asking now is: “What new exposure are we creating by integrating AI into our applications and business processes?” The urgency of that question grows as more departments embrace AI for customer support, analytics, software delivery, and internal automation. Every new use case often brings new plugins, new sources of data, and new identities, all of which can be weaponized by an adversary if not tightly governed and continuously tested.

Overview: The Shift in Enterprise Risk

The traditional "castle-and-moat" framework relied on a simple premise: build firewalls around trusted networks, segment internal systems, and control the few doorways leading into the environment. Software inside that perimeter operated on predictable logic. When a user issued a structured database query or an API call, the application executed hardcoded code to return expected data. Monitoring and detection programs could pattern-match on known behaviors, and change management moved at a slower pace, leaving time to validate each new integration.

AI integration dismantles that predictable environment. Connecting a Large Language Model (LLM) to internal databases, APIs, and automated workflows replaces predictable software logic with an unpredictable execution engine. An AI attack surface differs from a traditional software attack surface because it includes probabilistic model behavior, untrusted natural language inputs, and autonomous tool use. This increases the AI attack surface because an input submitted at the edge of a network can trigger a cascade of actions deep within internal infrastructure, bypassing perimeter checks entirely.

Understanding the exposure requires looking beyond the model’s responses. The risk grows when the model can use its permissions and integrations to retrieve data or take action across the surrounding application, i.e., an intelligence layer.

Attackers manipulating a customer-facing chatbot can end up navigating internal infrastructure faster than security operations teams can react. If security testing focuses exclusively on whether a model generates toxic language or bypasses basic prompt guardrails, critical vulnerabilities across the rest of the application stack remain completely invisible. Mature security programs must therefore extend their testing, detection, and governance frameworks to encompass the complete lifecycle of AI interactions, including upstream data ingestion and downstream action execution.

The Full-Stack AI Attack Surface

Evaluating risk accurately requires security teams to inspect the entire system architecture surrounding the intelligence layer. Model safety checks answer whether an LLM will output a forbidden recipe or reveal its system prompt, whereas system security assessments answer whether an attacker can trick that same model into wiping a production database or dumping client records.

An enterprise AI deployment consists of six intertwined layers, each presenting unique opportunities for compromise:

  1. The Model: The foundational LLM or fine-tuned model handling text generation, reasoning, and intent parsing. Attack surface at this layer includes prompt injection attacks, model inversion, jailbreaking, and training data poisoning.
  2. Agents and Tools: Autonomous logic, scripts, and plugin integrations that allow the model to execute real-world actions, run code, or trigger system processes. When a model receives agentic capabilities, a prompt injection vulnerability escalates from a text-display issue to an arbitrary execution issue. This is why agentic AI security must account for tool permissions and guardrails. Effective agentic AI security also restricts what tools can be invoked and with which credentials.
  3. Data and RAG Pipelines: Retrieval-Augmented Generation supplies an LLM with enterprise data so the model can answer with current context. Training data pipelines play a central role in the AI attack surface: poisoned documents, vector database pollution, and weak tenant isolation can leak or corrupt answers. Because model inputs include retrieved passages, attackers can seed malicious content that expands the AI attack surface via indirect prompt injection attacks inside the corpus.
  4. Application and APIs: The front-end interfaces, orchestration layers, and REST endpoints connecting users, the model, and internal microservices. Classic web flaws like Cross-Site Scripting (XSS), Server-Side Request Forgery (SSRF), and broken object-level authorization frequently re-emerge here, wrapped in natural language processing.
  5. Identity and Access Management: Service accounts, API keys, user tokens, and permission structures governing what the system can reach. Applications that run all AI-initiated actions through a single, over-privileged administrative service account flatten the internal security model.
  6. Underlying Infrastructure: Cloud environments, container registries, GPUs, orchestration engines, and storage buckets hosting the application stack. Misconfigurations in cluster management, unpatched hosting servers, or exposed management ports grant attackers raw access to underlying compute and memory.

Evaluating prompt safety while ignoring the application and identity layers is equivalent to locking a vault door while leaving the back window wide open.


Why Isolated AI Scanners Fall Short

Traditional application security tools and pure AI safety scanners take opposite (and equally incomplete) approaches:

  • Standard AppSec tools like static code analyzers and dynamic web scanners look for predictable software flaws but fail to comprehend natural language execution. A standard web scanner is unable to understand how natural language inputs translate into back-end operational commands. It cannot construct the complex, multi-turn conversational inputs required to guide an LLM into abusing an internal tool.
  • Pure AI safety scanners sit at the other extreme, testing whether prompt outputs violate corporate policy without looking at back-end execution. A scanner might fire hundreds of known jailbreak prompts at a model, observe that the model refused to output harmful text, and issue a clean report. That same scanner remains blind to the fact that the underlying RAG pipeline leaks confidential corporate records via simple, unauthenticated API requests.

Neither isolated approach captures the full attack surface. Effective evaluation pairs AI-specific prompt logic with standard application, API, cloud, and identity testing techniques.


How Weaknesses Chain Across the Stack

Security breaches in modern AI applications rarely happen in a single step. Risk stems from exploit chains, where a minor weakness in prompt handling can trigger a failure in back-end authorization or infrastructure.

To understand how an attack path unfolds across integrated boundaries, consider an exploit chain uncovered during technical testing of a cloud-native conversational AI application:

Figure 1: Exploit Chain
Figure 1: Exploit Chain
  1. Initial Manipulation: An attacker issues an indirect prompt injection to a user-facing chatbot, tricking the system into executing arbitrary code rather than processing text.
  2. Container Execution: The manipulated model invokes an internal code-execution tool, spawning an ephemeral microVM container within the cloud environment.
  3. Credential Exfiltration: The attacker uses the newly created execution environment to extract the underlying server credentials assigned to the container.
  4. Cross-Tenant Data Exposure: Because the service account possessed broad infrastructure permissions, the attacker uses the stolen credentials to bypass tenant isolation boundaries, reaching sensitive data stored across client organizations.

Prompt injection served as the initial spark. The real business impact occurred because of over-privileged service accounts, loose container boundaries, and weak cloud identity controls. Focusing purely on model behavior would have left the high-severity identity and infrastructure flaws completely unaddressed.


Managing Exposure and Testing

Inventory and governance should track capabilities, trust relationships, and access, not just which vendor model is used:

  • Capabilities: What actions can the system trigger autonomously? Can it send emails, write files, alter database rows, or execute system commands?
  • Trust Relationships: Which internal microservices, databases, and APIs trust the outputs of the AI application without secondary human or programmatic validation?
  • Access: What data stores, internal endpoints, and cloud credentials can the AI execution environment reach if an attacker compromises its logic?

A low-capability model with broad database permissions presents higher risk than an isolated, highly capable model.

Testing must follow the adversary path across layers. Pair model-focused checks with traditional AppSec, cloud, and identity reviews. AI pen testing should emulate realistic exploit chains, from inputs and RAG to tools and IAM, to surface the true business risk. Program leaders should adopt AI pen testing as a recurring control and ensure findings feed back into hardening data pipelines and tool permissions.

Take the Next Step in Securing Your AI Architecture:

  • Map Your AI Agents: Attackers can already find, connect to, and probe your exposed AI agent infrastructure. AIMap gives you that same visibility. Built by Bishop Fox, this open-source tool discovers, scores, and tests exposed AI endpoints so you can understand your real attack surface before someone else does.
  • Test Your AI Systems: Understanding your exposure is essential to building secure and resilient AI systems. Bishop Fox AI/LLM security assessments provide the experience and expertise to help you navigate this emerging threat landscape.

Banksy Fox exploder1

By Bishop Fox Researchers

Security Researchers

Due to the nature in which we conduct research and penetration tests, some of our security experts prefer to remain anonymous. Their work is published under our Bishop Fox name.

Bishop Fox is the leading authority in offensive security, providing solutions ranging from continuous penetration testing, red teaming, and attack surface management to product, cloud, and application security assessments. We’ve worked with more than 25% of the Fortune 100, half of the Fortune 10, eight of the top 10 global technology companies, and all of the top global media companies to improve their security. Our Cosmos platform, service innovation, and culture of excellence continue to gather accolades from industry award programs including Fast Company, Inc., SC Media, and others, and our offerings are consistently ranked as “world class” in customer experience surveys. We’ve been actively contributing to and supporting the security community for almost two decades and have published more than 16 open-source tools and 50 security advisories in the last five years. Learn more at bishopfox.com or follow us on Twitter.

Subscribe to our blog

Be first to learn about latest tools, advisories, and findings.