Build a Secure File-Import Service
Apply DARE end-to-end in your own repo to build a service that imports files from user-supplied URLs — hardened against SSRF, with content validation, asynchronous job processing, and an audit trail. Pulls the SSRF lab and the background-jobs playbook.
What you're building
"Import a file from a link" is one of the most common — and most dangerous —
features on the web. A user pastes a URL (a CSV to import, an avatar image, a
logo) and your server fetches it. The moment your server makes that request,
it does so from inside your infrastructure, carrying your network position
and your trust. A naive fetch(file_url) is a textbook Server-Side Request
Forgery (SSRF): the attacker never touches your internal network directly —
they hand your server a URL and let it do the reaching for them.
In this capstone you build a real file-import service end to end, in your
own repository, applying the full DARE method: POST /imports { file_url }
accepts a link, a safe-fetch layer refuses anything dangerous, a
background job downloads and validates the content asynchronously, the
result is persisted to storage, and every meaningful event is written to an
audit trail. This is not a starter you clone — it's a guide you apply. You
own the stack, the language, and the framework; DARE gives you the shape.
What you deliver
- An HTTP endpoint (
POST /imports) that accepts afile_url, enqueues work,
and returns immediately — never fetching inline in the web request. - A safe-fetch module that enforces, in order: HTTPS-only + URL validation,
DNS resolution with private/loopback/link-local blocking (including cloud
metadata169.254.169.254), per-hop redirect revalidation, and hard limits
on size, timeout, and content-type. - A worker/job that runs the safe-fetch, validates the downloaded content,
stores the result, and records the outcome — with failure handling and
idempotent retries. - An audit trail (an
importsrecord plus animport_eventstimeline)
that lets an operator answer "who imported what, from where, and what
happened" without leaking internal detail to the caller. - A test suite proving each defense fails closed and the happy path works.
How DARE guides you
You'll move through the method's phases, each as a milestone:
Design → the problem + threat model (SSRF, malicious files, exhaustion)
Blueprint → the endpoint contract, the SSRF validator rules, sync-vs-job,
the audit schema, storage
Tasks → decompose into a DAG with minimal dependencies
Execute → (a) the SSRF-safe fetch, then (b) the async import + audit
Harden → tests that prove every block, plus the submission bar
Two companion resources feed this project directly. The "Secure a Download
Endpoint Against SSRF" lab is the concept-first drill for the safe-fetch
pipeline you'll build in milestone 4. The "Run and Diagnose Background Jobs"
playbook is your reference for the worker, retries, and observability you'll
need in milestone 5. Keep both open as you go.
This is a guide, not code. Every snippet is pseudocode or a schema sketch —
translate it to your HTTP client, DNS resolver, job system, and test tools.
The goal is that your repo ends the project with a working, tested,
audited, SSRF-safe import service.
Architecture
Arquitetura de referência
O serviço é um endpoint de escrita fino na frente de um pipeline assíncrono. A
requisição web faz quase nada: valida o formato da entrada, cria um registro
import num estado pending, enfileira um job e retorna um id. Todo o trabalho
arriscado — DNS, conectar, baixar, validar conteúdo — acontece no worker, fora
da thread da requisição.
flowchart TD
U[Cliente] -->|POST /imports file_url| API[Endpoint de import]
API -->|checa formato + cria registro| DB[(tabela imports)]
API -->|enfileira| Q[[Fila de jobs]]
API -->|202 + id do import| U
Q --> W[Worker de import]
W --> SF{Guarda safe-fetch}
SF -->|1 só https + validação de URL| SF
SF -->|2 resolve DNS + bloqueia privado/loopback/link-local| SF
SF -->|3 revalida redirect por salto max 2| SF
SF -->|4 limites de tamanho + timeout + content-type| SF
SF -->|rejeita| AUD[(auditoria import_events)]
SF -->|bytes ok| CV[Validador de conteúdo]
CV -->|inválido| AUD
CV -->|válido| ST[(Object storage)]
ST --> DB
W -->|cada transição| AUD
BLK[[metadata 169.254.169.254 / 10.x / 127.x / fe80::]]:::danger -.->|bloqueado no passo 2| SF
classDef danger fill:#3b0d0d,stroke:#b71c1c,color:#fff;
Componentes
-
Endpoint de import (
POST /imports). Aceita{ file_url }. Valida só o
formato da requisição (presente, é string, dentro de um limite de tamanho),
escreve uma linhaimport(status: pending), enfileira o job e retorna
202com o id do import. Ele não busca nada. Quem chamou consulta
GET /imports/:idpara o status. -
Guarda safe-fetch. O firewall de SSRF, aplicado em ordem estrita: (1) faça
o parse da URL e exija esquemahttps, rejeiteuserinfoe esquemas
não-HTTPS; (2) resolva o host via DNS e rejeite se qualquer IP resolvido
for privado, loopback ou link-local (bloqueie127.0.0.0/8,10.0.0.0/8,
172.16.0.0/12,192.168.0.0/16,169.254.0.0/16— que contém o metadata da
cloud169.254.169.254—0.0.0.0/8,::1/128,fc00::/7,fe80::/10);
(3) desligue os auto-redirects e revalide o host a cada salto, com limite
de ~2; (4) imponha um timeout total, uma contagem máxima de bytes (header
e stream) e uma allowlist de content-type. -
Validador de conteúdo. Além dos limites de transporte, confirme que os
bytes são o que você pediu: farejar o tipo real (magic bytes, não só o
header) e, para imports estruturados (CSV), limitar contagem de linhas/colunas
e rejeitar em falha de parse. -
Object storage. Os bytes validados vão para object storage (S3/GCS/local),
indexados pelo id do import. O banco guarda um ponteiro, não o blob. -
Trilha de auditoria. Duas tabelas.
importsé o agregado (uma linha por
requisição, status atual).import_eventsé a linha do tempo append-only
(uma linha por transição:queued,fetch_rejected,content_rejected,
stored,failed,retried). A auditoria registra por que algo foi
rejeitado para os operadores; a API retorna só um status genérico a quem
chamou.
Esboço do modelo de dados
imports
id, user_id, file_url, status(pending|processing|succeeded|failed),
content_type, byte_size, storage_key, attempts, created_at, updated_at
import_events
id, import_id (fk), kind, detail(jsonb), created_at
# kind ∈ {queued, fetch_started, fetch_rejected, content_rejected,
# stored, failed, retried}
# detail guarda o motivo no servidor (IP bloqueado, tamanho, tipo) — nunca retornado ao cliente
Trade-offs para decidir no Blueprint
-
Síncrono vs. job. Buscar inline na requisição é mais simples mas errado
aqui: trava as threads web (DoS) e entrega ao atacante um oráculo de timing e
erro para mapear sua rede interna a partir das diferenças de resposta. O
import precisa rodar num job em background; o endpoint retorna antes de a
busca começar. Este é o padrão inegociável deste projeto. -
Onde validar. Divida: o endpoint valida só o formato (barato, síncrono,
rejeita lixo óbvio rápido). O worker valida destino e conteúdo (DNS, faixas
de IP, bytes) — as checagens caras e críticas de segurança que não podem rodar
na thread da requisição. -
Allowlist vs. denylist. Para esquemas e content-types, use
allowlist (sóhttps; os MIME types exatos que você aceita) — uma denylist
que você precisa manter completa é um jogo perdido. Para faixas de IP,
você está negando as faixas conhecidamente internas sobre o endereço
resolvido; se o seu caso permitir, aperte ainda mais para uma allowlist de
hosts de destino aprovados. Prefira a regra mais estreita que o seu produto
tolerar.
Milestones
-
Design — the problem and the threat model
Before a line of code, write the Design: what the service does, who uses
it, and — most importantly — how it can be abused. Importing a file from a
user-supplied URL is a request your server makes on the user's behalf from
inside your perimeter, so the threat model is the heart of this project.Produce a short design document that answers:
-
The feature. One paragraph: users submit a URL, the service imports the
file (CSV/image) asynchronously and stores the result; users can check
status. Name the concrete use cases you'll support. -
The actors. Who calls
POST /imports(authenticated users), who runs
the workers (your infra), and who the adversary is (any user who can put
a string infile_url). -
The threat model. Enumerate the abuse cases explicitly:
-
SSRF — the URL points at internal targets: cloud metadata
(169.254.169.254), loopback services (127.0.0.1:6379Redis), private
ranges (10.x,192.168.x), or non-HTTP schemes (file://,gopher://). -
Malicious files — a valid public URL that returns a
decompression bomb, a wrong/mislabeled content-type, or a file crafted to
break the downstream parser. -
Resource exhaustion — a huge file, a slow trickle (slow-loris), or a
flood of import requests, any of which starves memory, threads, or time. -
Information leakage — echoing upstream errors/timing back to the
caller, turning the endpoint into an oracle that maps your network.
-
SSRF — the URL points at internal targets: cloud metadata
-
Scope. What's in (safe fetch, async import, content validation, audit)
and explicitly out (e.g. virus scanning, user-facing retry UI) for this
capstone.
Done when: a
DESIGNdocument exists in your repo naming the feature,
the actors, and the four threat classes above (SSRF, malicious files,
exhaustion, leakage), each with at least one concrete attack example, plus an
explicit in/out scope list. If you can't name the attack, you can't defend
against it — this milestone is the source of truth for every later test. -
The feature. One paragraph: users submit a URL, the service imports the
-
Blueprint — endpoint contract, SSRF rules, storage & audit schema
Turn the Design into an architecture blueprint: the concrete contracts,
rules, and schema you'll build against. This is where you make the decisions
so that Execute is mechanical.Specify each of these:
-
Endpoint contract.
The write returns immediately; status is polled. The body carries noPOST /imports body: { "file_url": "https://cdn.example.com/data.csv" } 202 Accepted → { "id": "imp_123", "status": "pending" } GET /imports/:id 200 OK → { "id": "imp_123", "status": "succeeded|processing|failed" }
upstream detail on failure — only a generic status. -
SSRF validator rules (enumerated). Write them as an ordered, testable
list, because order is a security property:- Parse with a real URL parser; require
scheme == https; reject
userinfo; require a non-empty host. - Resolve the host to IPs; reject if any IP is in a blocked range
(127.0.0.0/8,10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,
169.254.0.0/16incl.169.254.169.254,0.0.0.0/8,::1/128,
fc00::/7,fe80::/10); use numeric IP-in-CIDR, cover IPv4-mapped IPv6. - Disable auto-redirects; on each hop re-run rules 1–2 on the new URL;
cap at ~2 hops. - Enforce total timeout, max bytes (Content-Length and streamed),
content-type allowlist.
- Parse with a real URL parser; require
-
Sync vs. job decision. State the rule and why: the fetch runs in a
background job, never inline — because inline fetches leak a
timing/error oracle and expose web threads to DoS. Define the queue, the
job name, and what the endpoint does synchronously (shape check + enqueue). -
Audit schema. Define the two tables from the architecture:
imports
(aggregate + current status) andimport_events(append-only timeline with
kindand a server-onlydetail). List the exact event kinds. -
Storage. Decide where validated bytes live (object storage keyed by
import id) and that the DB stores a pointer + metadata, not the blob.
Done when: a
BLUEPRINTdocument pins down the request/response shapes,
the four ordered SSRF rules as a numbered list, the explicit sync-vs-job
decision with rationale, theimports+import_eventsschema with named
event kinds, and the storage choice. Every later task should be able to cite
a line of this blueprint. -
Endpoint contract.
-
Tasks — decompose into a DAG with minimal dependencies
Break the blueprint into atomic, independently testable tasks and wire
their dependencies into a DAG. The goal is a graph where anything that can
be built in parallel is, and each edge is a real "needs the output of"
relationship — not an accident of how you happened to think about it.Suggested decomposition (adapt to your stack):
T1 URL validator (scheme=https, no userinfo, non-empty host) T2 DNS + IP guard (resolve host, block private/loopback/link-local) T3 Guarded fetch (T1 + T2 + no auto-redirect, per-hop revalidate, limits) T4 Import endpoint (shape check, create record, enqueue) — needs schema T5 Audit schema + writer (imports + import_events tables, event writer) T6 Import worker (run T3, validate content, store, audit via T5) T7 Retry + idempotency (safe re-run of T6 on failure) T8 Test suite (proves every block from T1–T7 fails closed)Then draw the dependency edges and keep them minimal:
flowchart LR T1[URL validator] --> T3[Guarded fetch] T2[DNS + IP guard] --> T3 T5[Audit schema + writer] --> T4[Import endpoint] T5 --> T6[Import worker] T3 --> T6 T4 --> T6 T6 --> T7[Retry + idempotency] T7 --> T8[Test suite] T3 --> T8Notice the parallelism the graph exposes: T1, T2, and T5 have no
dependencies and can be built at the same time. Keep the edges honest — if
a task doesn't truly need another's output, don't add the edge, or you'll
serialize work for no reason.Done when: you have a task list where each task is small enough to build
and test on its own, a DAG (dare-dag.yamlor equivalent) that validates
with no cycles and no broken references, and the graph makes the independent
tasks (validator, IP guard, audit) visibly parallel. Each task should name
its "done" check. -
Execute: SSRF validation — build the safe-fetch
Implement the safe-fetch — the SSRF firewall from tasks T1–T3. This is
the security core of the project; build it standalone and test it hard before
wiring it into the worker. The companion SSRF lab walks the concepts;
here you make it real in your repo.Build the pipeline in strict order — order is the security property:
MAX_REDIRECTS = 2 BLOCKED = [ "127.0.0.0/8","10.0.0.0/8","172.16.0.0/12","192.168.0.0/16", "169.254.0.0/16","0.0.0.0/8","::1/128","fc00::/7","fe80::/10" ] def safe_fetch(raw_url): url = raw_url for hop in 0..MAX_REDIRECTS: u = parse(url) # real parser, not regex if u.scheme != "https": reject # kills http/file/gopher/ftp/data if u.userinfo present: reject # https://ok@169.254.169.254/ if u.host empty: reject ips = dns_resolve(u.host) # may be several A/AAAA if ips empty: reject for ip in ips: # check EVERY resolved IP if numeric_in_cidr(ip, BLOCKED): reject # incl. 169.254.169.254 resp = http.get(u, follow_redirects=false, connect_to=ips, timeout=30s) if resp.is_redirect: url = resolve_relative(u, resp.header["Location"]) continue # re-validate the NEW url return read_limited(resp) # size + content-type limits reject("too many redirects")Non-negotiable details:
-
HTTPS-only, real parser, reject
userinfo— an allowlist of one scheme
beats a denylist you must keep complete;userinforemoves the
trusted.com@internalambiguity. -
Block on the resolved IP, numerically — parse each IP to bytes and test
CIDR containment. Immune to0x7f.0.0.1, decimal2130706433, and
zero-padding. Cover IPv4-mapped IPv6 (::ffff:127.0.0.1). Reject if any
resolved IP is internal. The169.254.0.0/16block (cloud metadata) is
mandatory. -
Drive redirects yourself, revalidate every hop — turn off auto-follow;
a302 → http://169.254.169.254/must hit the full gauntlet again. Cap
hops at ~2. -
Limits — total timeout; max bytes checked on the header and by
counting streamed bytes (Content-Length can lie); content-type allowlist;
cap decompressed size to blunt decompression bombs. -
Fail closed and quiet — every rejection returns a generic error; the
reason is logged server-side only.
Done when:
safe_fetchexists as an isolated, unit-tested module that,
with the network mocked, rejects: non-HTTPS schemes, a host resolving to
169.254.169.254/10.0.0.5/127.0.0.1, a redirect to an internal IP,
an oversized body (header and streamed), and a disallowed content-type — and
allows a well-formed public HTTPS file. It never echoes upstream detail. -
HTTPS-only, real parser, reject
-
Execute: async import — job, content validation, persistence & audit
Wire the safe-fetch into the asynchronous import pipeline (tasks T4–T7).
The endpoint enqueues; the worker does the risky work; every step is audited.
The companion "Run and Diagnose Background Jobs" playbook is your
reference for the worker, retries, and observability.Build it in two halves:
1. The endpoint (synchronous, tiny).
POST /imports (file_url): validate_shape(file_url) # present, string, length bound only imp = imports.create(user_id, file_url, status: "pending") audit(imp, kind: "queued") ImportJob.enqueue(imp.id) # returns before any fetch return 202, { id: imp.id, status: "pending" }2. The worker (asynchronous, guarded, audited).
ImportJob(import_id): imp = imports.find(import_id) imp.update(status: "processing"); audit(imp, "fetch_started") try: bytes = safe_fetch(imp.file_url) # milestone 4; may reject except Rejected as e: imp.update(status: "failed"); audit(imp, "fetch_rejected", detail: e.reason) return # do NOT retry a policy rejection if not valid_content(bytes, expected_type): # magic-byte sniff, CSV parse, row cap imp.update(status: "failed"); audit(imp, "content_rejected", detail: why) return key = storage.put(import_id, bytes) imp.update(status: "succeeded", storage_key: key, content_type: sniffed, byte_size: len(bytes)) audit(imp, "stored")Key decisions to get right:
-
Content validation is separate from transport limits. The safe-fetch
bounds size/type at the wire; here you confirm the bytes are what you
asked for — sniff magic bytes (don't trust the header), parse CSV and cap
rows/columns, reject on parse failure. -
Failure vs. retry. Distinguish a policy rejection (SSRF block, bad
content — deterministic, will fail again → do not retry, markfailed)
from a transient error (DNS blip, upstream 503, timeout → retry with
backoff). Only retry the transient class. -
Idempotent retries. A retried job must not double-store or double-audit.
Key storage byimport_id, make the "succeeded" transition a guarded
state change, and record aretriedevent so the timeline stays truthful. -
Audit every transition.
queued → fetch_started → (fetch_rejected | content_rejected | stored | failed | retried).detailholds the
server-only reason; the API still returns only a generic status.
Done when:
POST /importsenqueues and returns202without fetching;
the worker runs the safe-fetch, validates content, stores the result, and
writes an audit event for every transition; a policy rejection marksfailed
without retry while a transient error retries with backoff; and a retried job
is idempotent (no duplicate storage or audit rows). -
Content validation is separate from transport limits. The safe-fetch
-
Harden & verify — tests and submission criteria
A security control you can't test will rot. Close the project by proving,
with the network mocked, that every defense fails closed and the happy
path works — then meet the submission bar.Cover, at minimum, these cases (mocking lets you simulate dangerous responses
without real malicious infrastructure):test "rejects non-https scheme": expect_reject(safe_fetch("http://example.com/x")) expect_reject(safe_fetch("file:///etc/passwd")) test "rejects host resolving to a private / link-local IP": stub_dns("evil.test" => ["169.254.169.254"]) # cloud metadata expect_reject(safe_fetch("https://evil.test/")) stub_dns("lan.test" => ["10.0.0.5"]) expect_reject(safe_fetch("https://lan.test/")) test "rejects a redirect pointing at an internal IP": stub_http("https://ok.test/" => redirect_to("http://169.254.169.254/")) expect_reject(safe_fetch("https://ok.test/")) test "rejects an oversized file (header and streamed)": stub_http("https://big.test/" => body_of(500 MB)) expect_reject(safe_fetch("https://big.test/")) test "the import job processes a valid file and audits it": stub_dns("cdn.test" => ["93.184.216.34"]) # public IP stub_http("https://cdn.test/a.csv" => csv_response(1 MB)) imp = run_import("https://cdn.test/a.csv") expect(imp.status) == "succeeded" expect(audit_kinds(imp)) includes ["queued","fetch_started","stored"] test "a retry is idempotent (no duplicate storage or audit)": first = run_import_failing_once("https://cdn.test/a.csv") # transient error, then ok expect(imp.status) == "succeeded" expect(storage_objects(imp).count) == 1 expect(audit_kinds(imp)) includes ["retried"] test "never leaks the upstream error to the caller": stub_http("https://ok.test/" => connection_refused()) resp = post_imports("https://ok.test/") expect(resp.body) == generic_status() # no timing/error oracleAlso assert the cross-cutting rules:
POST /importsreturns before any fetch
(the work is in a job), and the redirect hop cap is enforced.
Submission criteria
Submit when all of the following hold in your repo:
-
Endpoint —
POST /imports { file_url }validates shape only, creates a
pendingimport, enqueues a background job, and returns202with an
id;GET /imports/:idreports status with no upstream detail. -
Safe-fetch — enforces, in order: HTTPS-only + URL validation → DNS
resolution blocking private/loopback/link-local (including
169.254.169.254) → per-hop redirect revalidation (cap ~2) → size +
timeout + content-type limits. -
Async import + audit — the worker runs the safe-fetch, validates
content, stores the result, and writes animport_eventsrow for every
transition; policy rejections markfailedwithout retry, transient
errors retry with backoff, and retries are idempotent. -
Tests (network mocked) proving each block: non-HTTPS, private/link-local
IP (incl. metadata), malicious redirect, oversized file — plus a happy-path
import that audits, an idempotent retry, and an "error is not leaked" test.
Include a short note on DNS rebinding hardening (connect-by-IP with a
pinnedHost) as a future step, and confirm no upstream error, body, or
timing is ever returned to the caller. -
Endpoint —