Attack LLMs and agents the way real adversaries do, then prove your defenses hold, across three hands-on days.
Classes are custom built from the following learning modules. (Times are approximate.) Custom formats from 2 to 5 days available.
Fingerprinting AI systems from the outside: identifying model providers, inferring model versions from response patterns, enumerating available tools and capabilities, and mapping the AI attack surface before you touch a prompt.
Building the attack plan. What are the high-value targets (training data, PII, internal tools, payment functions, admin capabilities)? What trust boundaries exist? Where are the weakest points? We build a threat model that drives the engagement, not a generic checklist.
Scope, authorization, data handling and the ethical boundaries of AI red teaming. What you can and cannot do with extracted data. Responsible handling of jailbreak techniques.
Hands on: the core injection techniques. Role override, instruction smuggling, delimiter escape, encoding bypass (base64, rot13, unicode tricks) and multi-turn state manipulation. We attack a live target and observe what works, what fails and why.
Hands on: inject payloads through data the model retrieves rather than through the user prompt. Poisoned documents in RAG, malicious tool responses, hidden instructions in web pages and email content. The attack surface where the user never typed the injection.
Payload obfuscation, token-level attacks, few-shot poisoning, system prompt extraction and context window manipulation. The techniques that bypass first-generation defenses.
The categories of jailbreak that work today: roleplay (DAN, character assumption), hypothetical framing, encoding and language switching, multi-turn escalation and social engineering the model. We test each category against live targets.
Hands on: systematically test guardrails rather than throwing random prompts. Build a test matrix: categories of restricted content, categories of evasion technique, and the cross-product that tells you which guardrails hold and which do not. Produce a coverage report.
Hands on: configure and run promptfoo against a live target. Write test cases, define success criteria, run the suite and interpret the results.
Hands on: attempt to extract training data from a model. Verbatim memorization probes, membership inference and canary detection. What the model remembers and how to surface it.
Hands on: extract the system prompt from a deployed model. Direct extraction, indirect inference and the side channels that leak system instructions even when the model is told not to reveal them.
Querying a model systematically to build a functional copy. Distillation attacks, output harvesting and what rate limiting and output controls actually prevent.
Crafting inputs that cause the model to misclassify, misinterpret or produce incorrect output. Adversarial suffixes, perturbation attacks and the gap between academic adversarial ML and practical red teaming.
Hands on: escalate an agent's effective privileges. Manipulate the agent into calling tools it should not call, accessing data outside its scope, or performing actions above its trust level. Confused deputy attacks in an agent context.
Hands on: poison tool descriptions to make the agent do something the tool owner controls. Change tool behaviour between listing and invocation. Inject through tool responses.
When agents talk to other agents, the attack surface multiplies. Compromise one agent and use it to manipulate others through the delegation chain. Scope escalation across agent boundaries.
Hands on: attack MCP servers using the techniques from the CVE case studies. Replay, DoS (CVE-2026-39313), unauthenticated enumeration, tool rug-pulls and injection through tool responses. dvmcp
Hands on: advanced promptfoo usage. Custom providers, adversarial plugins, grading functions and CI integration. Building a promptfoo test suite that runs as part of your red team engagement.
Hands on: configure and run garak (the NVIDIA LLM vulnerability scanner) against a target. Understand its probe categories, interpret results and integrate findings into your report.
Hands on: Microsoft's Python Risk Identification Toolkit. Multi-turn attack orchestration, automated jailbreak chains and scored evaluation. When to use PyRIT vs promptfoo vs garak.
When the tools do not cover your target. Writing custom attack scripts, building prompt mutation engines and the discipline of logging every attempt for the report.
Finding structure: attack technique, reproduction steps, impact assessment, AISVS control mapping and OWASP LLM Top 10 mapping. Writing findings that defenders can fix, not just read.
Report structure for AI red team engagements. Executive summary, scope, methodology, findings by severity, and recommendations. What to include, what to redact and how to handle sensitive findings.
Working with the blue team and the AI engineering team after the engagement. Reproducing findings, verifying fixes and the re-test cycle.
Attack a complete AI system (LLM + RAG + agents + MCP tools) end to end. Reconnaissance, injection, jailbreak, data extraction, agent escalation and tool poisoning. Produce a scored findings report. dvmcp dvrag
Systematic guardrail testing against a hardened target. Build the test matrix, run automated and manual attacks, document which guardrails hold and which break.
Red-team an agent estate. Escalate privileges, manipulate delegation chains, exfiltrate data through tool calls and produce the evidence for each finding.
Hands-on with the CyberSecAI tooling, not slideware.
Jailbreak and injection fuzzer with a payload corpus and pass/fail scoring.
LLM forensic auditor used to strip guardrail layers and probe an unguarded model in the lab.
Static scanner running the OWASP-standards check set.
The vulnerable RAG and MCP targets you attack.
Add AI attack surface and adversarial testing to your remit.
Bring application security discipline to LLMs and agents.
See how your models fail under adversarial pressure, and fix it.
Move into AI targets with methods and tooling that carry over.
Sharpen a structured, repeatable practice with a full toolchain.
No slideware. You attack real LLMs and agents on live instances backed by cloud GPUs, reproduce the paths you find, and prove which defenses hold and which do not.
This course counts toward the CyberSecAI AISVS assurance track. Assessment is practical: you run a structured red-team against a live LLM or agent, reproduce the injection, jailbreak and tool-abuse paths you find, measure the adversarial-ML exposure, and write the report that engineering can act on.
Three days, hands-on, attacking live models and agents, taught by the former OWASP-AISVS Co-Leader (v1.0). Places are limited.
Corporate & sovereign cohorts, and bespoke on-site delivery, on request.