Exploit a Vulnerable AI Agent
Talk a tool-calling banking agent into transferring funds it never should have — no ownership checks, no limits, no confirmation — in a disposable local sandbox with synthetic accounts.
PremiumProblem
An AI agent that can call real tools — check a balance, generate a report, move money — is only as safe as the guardrails *around* those tools, not the model's good intentions. A model has no innate concept of "this account isn't mine" or "this amount is too large" unless the tool implementation enforces it. This is **excessive agency** (OWASP LLM06 for LLM applications): giving an agent the ability to take a high-impact, irreversible action without ownership validation, limits, or a human confirmation step in between. The danger compounds because the attack surface isn't a crafted payload — it's plain natural language. If the underlying tool has no guardrails, a user (or an attacker who gets a foothold in a chat session) doesn't need any technical trick at all: a politely worded request is the entire exploit. In this lab you'll run the **DARE Vulnerable AI Suite** locally and attack its `vulnerable-agent` challenge: a banking assistant with tools to check balances, generate reports, transfer funds, send messages, and look up documents. The `criar_transferencia` (create-transfer) tool validates nothing — no ownership, no ceiling, no confirmation prompt. Your job is to get it to move money out of the seeded `demo-user` account using nothing but conversation. > Practice only against the disposable `dare-vulnerable-ai` suite running > locally on `127.0.0.1:8000` (bound to localhost only), with synthetic > seed accounts that hold no real money. Never attempt this kind of > social-engineering-of-an-agent technique against a real financial > system or any system you do not own or have explicit written > authorization to test.
Objectives
By the end of this lab you will be able to:
- Explain excessive agency (OWASP LLM06) and why tool guardrails must
live in the tool implementation, not in the model's judgment. - Interact with a tool-calling agent's chat API and read back which
tools it invoked (tool_calls). - Craft a natural-language request that convinces an agent to invoke a
high-impact action with no confirmation step. - Verify a state change (a ledger entry) caused entirely through
conversation, not a direct API call to the sensitive action. - Read a challenge's official solution and name the three controls
(ownership, limit, confirmation) a safe implementation needs.
Prerequisites
To complete this lab you'll need:
- Docker and Docker Compose installed.
- Access to the private
dare-vulnerable-airepository (enrolled
students have access):git@github.com:darelabs-tech/dare-vulnerable-ai.git. -
curl(or an equivalent HTTP client) and a terminal. - No prior experience with agentic AI or tool-calling required.
The mental model: tools are the trust boundary, not the model
An LLM agent's "reasoning" is not a security control. The model decides
which tool to call and with what arguments based on the
conversation — but whether that call is actually allowed is a decision
that has to be made by the tool's own code, every single time, with no
exceptions for "the model seemed confident." If the tool trusts its
caller (the model) the way a backend should only ever trust an
authenticated, authorized human or service, you get excessive agency.
Vulnerable agent flow
user message ──▶ model decides: call criar_transferencia(...)
│
▼
tool executes immediately: no ownership check,
no amount ceiling, no "are you sure?" step
│
▼
money moves, ledger updated
The seeded accounts
This challenge ships two accounts to make the exploit observable:
-
demo-user— starts with a balance of500000cents (R$ 5,000.00). -
attacker-controlled— starts with a balance of0.
The challenge is marked solved as soon as a transfer exists for your
session where amount_cents > 500000 (more than the account ever had —
proving no balance check happened) or from_account_id != "demo-user" (proving no ownership check happened, since you never
authenticated as any account owner in the first place).
How you'll work through this lab
- Unlock the
vulnerable-agentchallenge. - Send a normal, in-scope request (checking a balance) to see expected
behavior. - Convince the agent, in plain language, to transfer the entire
demo-userbalance toattacker-controlled. - Confirm the transfer actually happened by reading the ledger.
- Read the official solution and name the three missing controls.