Companies are shipping chatbots, AI copilots, retrieval-augmented generation (RAG) search and autonomous agents into production at remarkable speed. These systems introduce risks that traditional application testing doesn't fully cover: an attacker no longer needs a code bug if they can simply talk the model into misbehaving. This guide explains what AI and LLM penetration testing involves, the OWASP framework behind it, and how to scope an assessment.
What is AI / LLM penetration testing?
It is a security assessment of applications that use large language models — not of the foundation model itself. Testers examine how your application combines prompts, user input, retrieved documents, tools and plugins, and whether an attacker can manipulate that combination to leak data, take unauthorised actions or cause harm. It is usually performed alongside a conventional web/API test of the same application, because most AI features sit on top of ordinary APIs.
The OWASP Top 10 for LLM Applications (2025)
| ID | Risk | What it means for your app |
|---|---|---|
| LLM01 | Prompt Injection | User input or retrieved content overrides your instructions — directly, or indirectly via documents, web pages or emails the model reads |
| LLM02 | Sensitive Information Disclosure | The model reveals personal data, secrets or other users' information |
| LLM03 | Supply Chain | Risks from third-party models, datasets, plugins and libraries |
| LLM04 | Data and Model Poisoning | Manipulated training, fine-tuning or embedding data changes behaviour |
| LLM05 | Improper Output Handling | Model output passed unsafely into browsers, databases, shells or other systems (e.g., XSS or injection via the model) |
| LLM06 | Excessive Agency | The model or agent has more tools, permissions or autonomy than it needs |
| LLM07 | System Prompt Leakage | Hidden instructions — sometimes containing secrets or logic — are revealed |
| LLM08 | Vector and Embedding Weaknesses | RAG stores leak data across users or tenants, or can be manipulated |
| LLM09 | Misinformation | Confident but false outputs that users or systems rely on |
| LLM10 | Unbounded Consumption | Abuse that drives runaway cost or denial of service |
What testers actually test
Prompt injection and jailbreaks
Testers attempt to override system instructions directly through chat input and indirectly by planting instructions in content the model will later read — an uploaded file, a support ticket, a web page, an email in a connected inbox. Indirect injection is often the more dangerous variant because the victim never sees the malicious instruction.
Data access and isolation
In RAG systems, the key question is whether retrieval respects the user's permissions. Testers check whether one user or tenant can get the model to surface another's documents, and whether sensitive data in the index can be extracted by clever questioning.
Tool use and agents
If the model can call APIs, send emails, run code, query databases or make purchases, testers try to make it do so on an attacker's behalf. The most important control is ensuring the tools themselves enforce the user's authorization — never relying on the model to decide.
Output handling
Testers check whether model output is rendered as HTML, used in SQL or shell commands, or fed into other systems without validation — classic injection risks with a new delivery path.
Abuse and cost
Rate limits, token limits and cost controls are tested to make sure a single user can't run up large bills or degrade service for everyone.
Design principles that reduce AI risk
- Treat all model input — including retrieved content — as untrusted
- Enforce authorization in tools and data layers, not in prompts
- Give agents the minimum tools and permissions needed; require human approval for high-impact actions
- Keep secrets out of system prompts
- Validate and encode model output before using it elsewhere
- Log prompts, tool calls and outputs for monitoring and investigation, with privacy controls
Scoping an AI security assessment
Share the architecture (model provider, orchestration framework, vector store, tools and connectors), user roles, data sources the model can read, actions it can take, and any guardrails or content filters. Most engagements combine 3–10 tester-days of AI-specific testing with a standard web/API test of the surrounding application.
Frameworks and regulation to be aware of
- OWASP Top 10 for LLM Applications and the broader OWASP GenAI Security Project
- MITRE ATLAS — adversary tactics and techniques against AI systems
- NIST AI Risk Management Framework
- ISO/IEC 42001 — AI management systems
- EU AI Act — risk-based obligations, including robustness and cybersecurity requirements for high-risk systems
Frequently asked questions
Can prompt injection be fully prevented?
Not reliably with today's models. The practical approach is to limit impact: treat model input as untrusted, restrict tools and permissions, and enforce authorization outside the model.
Do we need to test the AI model provider itself?
No. You test your application's use of the model — prompts, data access, tools and output handling. Model providers have their own security programs and terms.
Is AI testing a separate engagement?
It can be, but it is usually combined with a web/API test of the same application, since most AI features sit on ordinary APIs.
We offer dedicated AI & LLM security testing. Request a quote and select "AI / LLM application".