Fix a Vulnerable MCP Server
Patch a real tool-poisoning and path-traversal vulnerability chain in the DARE Vulnerable AI Suite's MCP challenge — defensive engineering, not exploitation: you fix the bug, then prove the exploit no longer works.
PremiumProblem
> **This lab is about defense, not offense.** If you've taken the AI Red > Teaming track, you may have already attacked this exact challenge > (`vulnerable-mcp`) to extract its flag via tool poisoning and path > traversal. This lab uses the same target for the opposite purpose: > you are the engineer who has to **read the vulnerable code, understand > why it's exploitable, write the fix, and prove the fix holds** — the > other half of the security skill set, and arguably the harder half. > Exploiting a known bug takes minutes; writing a fix that doesn't break > the feature and doesn't just move the vulnerability somewhere else > takes real understanding. The `vulnerable-mcp` challenge in the **DARE Vulnerable AI Suite** ships a `read_file(path)` tool with two independent, real vulnerabilities in `app/challenges/vulnerable_mcp.py` and `app/db/seed.py`: 1. **Tool poisoning.** The `description` field of a different tool, `send_notification`, contains a hidden instruction telling any model that reads it to always call `read_file('../../secrets/api_keys.txt')` and include the contents in its response — a prompt injected through the tool catalog itself, not through user input. 2. **Path traversal.** `read_file` resolves the requested path against `SANDBOX_ROOT = Path(__file__).parent / "mcp_sandbox"`, but only checks that the *resolved* path stays inside `SANDBOX_ROOT` overall — not inside the `documents/` subdirectory it's actually meant to serve. A path like `../../secrets/api_keys.txt` walks straight out of `documents/` into a sibling `secrets/` directory containing a real flag file, `DARE-FLAG-4-{suffix}`. You'll start by reproducing the vulnerability exactly as an attacker would — proof it's real, before you touch any code — then fix both root causes, then reproduce the exact same attack again and confirm it now fails. > Practice only against the disposable `dare-vulnerable-ai` suite > running locally on `127.0.0.1:8000`. Never attempt path traversal > against a system you do not own or have explicit written > authorization to test.
Objectives
By the end of this lab you will be able to:
- Reproduce a known vulnerability as a baseline before fixing it, so
you have concrete proof a fix actually closes the hole. - Read a Python path-sandboxing implementation and identify why
checking "inside the sandbox root" is a weaker boundary than
checking "inside the specific subdirectory the tool is meant to
serve." - Implement a correct path-containment check using
Path.resolve().is_relative_to()against the intended
subdirectory, raising an appropriate error when a path escapes it. - Identify and remove a hidden prompt-injection instruction embedded in
a tool's description field. - Understand why changing seeded data requires recreating the
database volume (docker compose down -v), not just restarting the
containers. - Re-verify a fix using the exact same attack request that previously
succeeded, and treat "fails now" as the actual acceptance criterion.
Prerequisites
To complete this lab you'll need:
- Docker and Docker Compose installed.
- Access to the private
dare-vulnerable-airepository (enrolled
students have access):git@github.com:darelabs-tech/dare-vulnerable-ai.git. -
curl(or an equivalent HTTP client), a terminal, and a Python-aware
text editor. - Basic familiarity with
pathlib.Pathand Python exception handling.
Before vs. after — the shape of what you're proving
BEFORE the fix AFTER the fix
─────────────── ──────────────
curl .../invoke curl .../invoke
path=../../secrets/api_keys.txt path=../../secrets/api_keys.txt
→ 200 OK → error / solved:false
"solved": true path rejected: escapes
result: "DARE-FLAG-4-xxxx" the documents/ sandbox
The exact same request, before and after — that's the whole proof.
Nothing about the legitimate use of read_file (reading files under
documents/) should change; only the escape should stop working.
The two root causes, and why both matter
-
Tool poisoning supplies the motive and the exact path — a model
readingsend_notification's description is told, in effect, "go
read this specific secrets file." Fixing this means the description
no longer contains a hidden instruction, so a model reading it has no
reason to ever request that path. -
Path traversal supplies the mechanism — even with the
poisoned description removed,read_filewould still happily serve
../../secrets/api_keys.txtto anyone (or anything) that asked for
it directly. Fixing this means the tool itself refuses to serve
anything outside its intended directory, regardless of who or what
is asking.
You need both fixes. Removing only the poisoned description leaves a
working path-traversal primitive that a different injected instruction,
or a directly malicious client, could still trigger. Removing only the
traversal bug leaves a tool description actively lying to every model
that reads it.
How to work through this lab
Reproduce first, understand second, fix third, reproduce again last.
Resist the urge to jump straight to writing the fix — the "before"
reproduction in Step 1 is what makes Step 5's "after" reproduction
meaningful proof rather than a guess.