24/7 SOC monitoring & incident responsesales@bugfoe.com
BestPentestingby BugFoe
Emerging risk

AI & LLM Application Penetration Testing

Chatbots, copilots, RAG and agents add risks traditional testing misses. Here's how they're tested — and how to design them safely.

Updated October 20264 min readBy the BestPentesting Security Research Team

Companies are shipping chatbots, AI copilots, retrieval-augmented generation (RAG) search and autonomous agents into production at remarkable speed. These systems introduce risks that traditional application testing doesn't fully cover: an attacker no longer needs a code bug if they can simply talk the model into misbehaving. This guide explains what AI and LLM penetration testing involves, the OWASP framework behind it, and how to scope an assessment.

What is AI / LLM penetration testing?

It is a security assessment of applications that use large language models — not of the foundation model itself. Testers examine how your application combines prompts, user input, retrieved documents, tools and plugins, and whether an attacker can manipulate that combination to leak data, take unauthorised actions or cause harm. It is usually performed alongside a conventional web/API test of the same application, because most AI features sit on top of ordinary APIs.

The OWASP Top 10 for LLM Applications (2025)

IDRiskWhat it means for your app
LLM01Prompt InjectionUser input or retrieved content overrides your instructions — directly, or indirectly via documents, web pages or emails the model reads
LLM02Sensitive Information DisclosureThe model reveals personal data, secrets or other users' information
LLM03Supply ChainRisks from third-party models, datasets, plugins and libraries
LLM04Data and Model PoisoningManipulated training, fine-tuning or embedding data changes behaviour
LLM05Improper Output HandlingModel output passed unsafely into browsers, databases, shells or other systems (e.g., XSS or injection via the model)
LLM06Excessive AgencyThe model or agent has more tools, permissions or autonomy than it needs
LLM07System Prompt LeakageHidden instructions — sometimes containing secrets or logic — are revealed
LLM08Vector and Embedding WeaknessesRAG stores leak data across users or tenants, or can be manipulated
LLM09MisinformationConfident but false outputs that users or systems rely on
LLM10Unbounded ConsumptionAbuse that drives runaway cost or denial of service

What testers actually test

Prompt injection and jailbreaks

Testers attempt to override system instructions directly through chat input and indirectly by planting instructions in content the model will later read — an uploaded file, a support ticket, a web page, an email in a connected inbox. Indirect injection is often the more dangerous variant because the victim never sees the malicious instruction.

Data access and isolation

In RAG systems, the key question is whether retrieval respects the user's permissions. Testers check whether one user or tenant can get the model to surface another's documents, and whether sensitive data in the index can be extracted by clever questioning.

Tool use and agents

If the model can call APIs, send emails, run code, query databases or make purchases, testers try to make it do so on an attacker's behalf. The most important control is ensuring the tools themselves enforce the user's authorization — never relying on the model to decide.

Output handling

Testers check whether model output is rendered as HTML, used in SQL or shell commands, or fed into other systems without validation — classic injection risks with a new delivery path.

Abuse and cost

Rate limits, token limits and cost controls are tested to make sure a single user can't run up large bills or degrade service for everyone.

Design principles that reduce AI risk

  • Treat all model input — including retrieved content — as untrusted
  • Enforce authorization in tools and data layers, not in prompts
  • Give agents the minimum tools and permissions needed; require human approval for high-impact actions
  • Keep secrets out of system prompts
  • Validate and encode model output before using it elsewhere
  • Log prompts, tool calls and outputs for monitoring and investigation, with privacy controls

Scoping an AI security assessment

Share the architecture (model provider, orchestration framework, vector store, tools and connectors), user roles, data sources the model can read, actions it can take, and any guardrails or content filters. Most engagements combine 3–10 tester-days of AI-specific testing with a standard web/API test of the surrounding application.

Frameworks and regulation to be aware of

  • OWASP Top 10 for LLM Applications and the broader OWASP GenAI Security Project
  • MITRE ATLAS — adversary tactics and techniques against AI systems
  • NIST AI Risk Management Framework
  • ISO/IEC 42001 — AI management systems
  • EU AI Act — risk-based obligations, including robustness and cybersecurity requirements for high-risk systems

Frequently asked questions

Can prompt injection be fully prevented?

Not reliably with today's models. The practical approach is to limit impact: treat model input as untrusted, restrict tools and permissions, and enforce authorization outside the model.

Do we need to test the AI model provider itself?

No. You test your application's use of the model — prompts, data access, tools and output handling. Model providers have their own security programs and terms.

Is AI testing a separate engagement?

It can be, but it is usually combined with a web/API test of the same application, since most AI features sit on ordinary APIs.

We offer dedicated AI & LLM security testing. Request a quote and select "AI / LLM application".

BestPentesting Security Research Team
Written and technically reviewed by practising penetration testers and SOC analysts. Last reviewed October 2026. See our methodology.

Ready to find your risks before attackers do?

Tell us what you need tested or monitored. A senior consultant replies within one business day with a scoped, fixed-price proposal.

  • Fixed-price proposal
  • Reply within 1 business day
  • NDA on request