LLM & RAG Red Teaming
By Wanderson Leandro de Oliveira
This course takes you from theory straight into exploitation: red teaming large language models and retrieval-augmented generation (RAG) pipelines, entirely against the DARE Vulnerable AI Suite running on your own machine. You will start with direct prompt injection — breaking a chatbot's system prompt with instruction override and extraction techniques — and then move into more advanced territory: indirect injection via documents and tool output, multi-turn payload splitting, encoding tricks, and why textual jailbreak guardrails are inherently fragile. From there the course turns to the model and data layer: training data extraction and memorization, data poisoning and backdoor concepts, and denial-of-wallet attacks that drain a target's token budget. The second half is dedicated to RAG: how ingestion, chunking, embeddings, retrieval and re-ranking actually work, where metadata filtering and tenant isolation break down, and how a poisoned document can leak a competitor tenant's secrets straight into an answer. Two lessons are fully hands-on against real, intentionally vulnerable challenges in the DARE Vulnerable AI Suite — an instruction-override chat endpoint and a multi-tenant RAG pipeline with no metadata filter — where you will capture real flags with curl and then read the exact fix. The course closes by teaching you to measure and communicate what you found: Attack Success Rate and other AI red teaming metrics, mapping findings to the OWASP LLM Top 10 and MITRE ATLAS, and writing an executive summary a non-technical stakeholder can act on.
Course content
Direct Prompt Injection
- 🔒 Instruction override: breaking a system prompt text
- 🔒 System prompt extraction techniques text
- 🔒 Hands-on: breaking the llm-chat challenge text
Advanced Prompt Attacks
- 🔒 Indirect prompt injection via documents and tools text
- 🔒 Multi-turn attacks, payload splitting and encoding text
- 🔒 Jailbreaks, roleplay bypass and why guardrails fail text
Model and Data Vulnerabilities
- 🔒 Training data extraction and memorization text
- 🔒 Data poisoning and backdoors: concepts text
- 🔒 Denial of wallet and resource exhaustion text
RAG Architecture and Attack Surface
- 🔒 RAG architecture: ingestion, chunking, retrieval text
- 🔒 Where RAG breaks: metadata filtering and tenant isolation text
- 🔒 RAG poisoning and document injection text
Exploiting a Vulnerable RAG Pipeline
- 🔒 Finding cross-tenant leakage in retrieval text
- 🔒 Hands-on: breaking the vulnerable-rag challenge text
- 🔒 Fixing it: authorization inside the vector store text
Measuring and Reporting Findings
- 🔒 Attack Success Rate and AI red teaming metrics text
- 🔒 Mapping findings to OWASP LLM Top 10 and MITRE ATLAS text
- 🔒 Writing an executive summary text