Security

Build a Secure File-Import Service

Apply DARE end-to-end in your own repo to build a service that imports files from user-supplied URLs — hardened against SSRF, with content validation, asynchronous job processing, and an audit trail. Pulls the SSRF lab and the background-jobs playbook.

What you're building

"Import a file from a link" is one of the most common — and most dangerous —
features on the web. A user pastes a URL (a CSV to import, an avatar image, a
logo) and your server fetches it. The moment your server makes that request,
it does so from inside your infrastructure, carrying your network position
and your trust. A naive fetch(file_url) is a textbook Server-Side Request
Forgery (SSRF)
: the attacker never touches your internal network directly —
they hand your server a URL and let it do the reaching for them.

In this capstone you build a real file-import service end to end, in your
own repository, applying the full DARE method: POST /imports { file_url }
accepts a link, a safe-fetch layer refuses anything dangerous, a
background job downloads and validates the content asynchronously, the
result is persisted to storage, and every meaningful event is written to an
audit trail. This is not a starter you clone — it's a guide you apply. You
own the stack, the language, and the framework; DARE gives you the shape.

What you deliver

How DARE guides you

You'll move through the method's phases, each as a milestone:

Design    → the problem + threat model (SSRF, malicious files, exhaustion)
Blueprint → the endpoint contract, the SSRF validator rules, sync-vs-job,
            the audit schema, storage
Tasks     → decompose into a DAG with minimal dependencies
Execute   → (a) the SSRF-safe fetch, then (b) the async import + audit
Harden    → tests that prove every block, plus the submission bar

Two companion resources feed this project directly. The "Secure a Download
Endpoint Against SSRF"
lab is the concept-first drill for the safe-fetch
pipeline you'll build in milestone 4. The "Run and Diagnose Background Jobs"
playbook is your reference for the worker, retries, and observability you'll
need in milestone 5. Keep both open as you go.

This is a guide, not code. Every snippet is pseudocode or a schema sketch —
translate it to your HTTP client, DNS resolver, job system, and test tools.
The goal is that your repo ends the project with a working, tested,
audited, SSRF-safe import service.

Architecture

Arquitetura de referência

O serviço é um endpoint de escrita fino na frente de um pipeline assíncrono. A
requisição web faz quase nada: valida o formato da entrada, cria um registro
import num estado pending, enfileira um job e retorna um id. Todo o trabalho
arriscado — DNS, conectar, baixar, validar conteúdo — acontece no worker, fora
da thread da requisição.

flowchart TD
    U[Cliente] -->|POST /imports file_url| API[Endpoint de import]
    API -->|checa formato + cria registro| DB[(tabela imports)]
    API -->|enfileira| Q[[Fila de jobs]]
    API -->|202 + id do import| U
    Q --> W[Worker de import]
    W --> SF{Guarda safe-fetch}
    SF -->|1 só https + validação de URL| SF
    SF -->|2 resolve DNS + bloqueia privado/loopback/link-local| SF
    SF -->|3 revalida redirect por salto max 2| SF
    SF -->|4 limites de tamanho + timeout + content-type| SF
    SF -->|rejeita| AUD[(auditoria import_events)]
    SF -->|bytes ok| CV[Validador de conteúdo]
    CV -->|inválido| AUD
    CV -->|válido| ST[(Object storage)]
    ST --> DB
    W -->|cada transição| AUD
    BLK[[metadata 169.254.169.254 / 10.x / 127.x / fe80::]]:::danger -.->|bloqueado no passo 2| SF
    classDef danger fill:#3b0d0d,stroke:#b71c1c,color:#fff;

Componentes

Esboço do modelo de dados

imports
  id, user_id, file_url, status(pending|processing|succeeded|failed),
  content_type, byte_size, storage_key, attempts, created_at, updated_at

import_events
  id, import_id (fk), kind, detail(jsonb), created_at
  # kind ∈ {queued, fetch_started, fetch_rejected, content_rejected,
  #         stored, failed, retried}
  # detail guarda o motivo no servidor (IP bloqueado, tamanho, tipo) — nunca retornado ao cliente

Trade-offs para decidir no Blueprint

Milestones

  1. Design — the problem and the threat model

    Before a line of code, write the Design: what the service does, who uses
    it, and — most importantly — how it can be abused. Importing a file from a
    user-supplied URL is a request your server makes on the user's behalf from
    inside your perimeter
    , so the threat model is the heart of this project.

    Produce a short design document that answers:

    • The feature. One paragraph: users submit a URL, the service imports the
      file (CSV/image) asynchronously and stores the result; users can check
      status. Name the concrete use cases you'll support.
    • The actors. Who calls POST /imports (authenticated users), who runs
      the workers (your infra), and who the adversary is (any user who can put
      a string in file_url).
    • The threat model. Enumerate the abuse cases explicitly:
      • SSRF — the URL points at internal targets: cloud metadata
        (169.254.169.254), loopback services (127.0.0.1:6379 Redis), private
        ranges (10.x, 192.168.x), or non-HTTP schemes (file://, gopher://).
      • Malicious files — a valid public URL that returns a
        decompression bomb, a wrong/mislabeled content-type, or a file crafted to
        break the downstream parser.
      • Resource exhaustion — a huge file, a slow trickle (slow-loris), or a
        flood of import requests, any of which starves memory, threads, or time.
      • Information leakage — echoing upstream errors/timing back to the
        caller, turning the endpoint into an oracle that maps your network.
    • Scope. What's in (safe fetch, async import, content validation, audit)
      and explicitly out (e.g. virus scanning, user-facing retry UI) for this
      capstone.

    Done when: a DESIGN document exists in your repo naming the feature,
    the actors, and the four threat classes above (SSRF, malicious files,
    exhaustion, leakage), each with at least one concrete attack example, plus an
    explicit in/out scope list. If you can't name the attack, you can't defend
    against it — this milestone is the source of truth for every later test.

  2. Blueprint — endpoint contract, SSRF rules, storage & audit schema

    Turn the Design into an architecture blueprint: the concrete contracts,
    rules, and schema you'll build against. This is where you make the decisions
    so that Execute is mechanical.

    Specify each of these:

    • Endpoint contract.
      POST /imports
      body: { "file_url": "https://cdn.example.com/data.csv" }
      202 Accepted → { "id": "imp_123", "status": "pending" }
      
      GET /imports/:id
      200 OK → { "id": "imp_123", "status": "succeeded|processing|failed" }
      
      The write returns immediately; status is polled. The body carries no
      upstream detail on failure — only a generic status.
    • SSRF validator rules (enumerated). Write them as an ordered, testable
      list, because order is a security property:
      1. Parse with a real URL parser; require scheme == https; reject
        userinfo; require a non-empty host.
      2. Resolve the host to IPs; reject if any IP is in a blocked range
        (127.0.0.0/8, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16,
        169.254.0.0/16 incl. 169.254.169.254, 0.0.0.0/8, ::1/128,
        fc00::/7, fe80::/10); use numeric IP-in-CIDR, cover IPv4-mapped IPv6.
      3. Disable auto-redirects; on each hop re-run rules 1–2 on the new URL;
        cap at ~2 hops.
      4. Enforce total timeout, max bytes (Content-Length and streamed),
        content-type allowlist.
    • Sync vs. job decision. State the rule and why: the fetch runs in a
      background job, never inline — because inline fetches leak a
      timing/error oracle and expose web threads to DoS. Define the queue, the
      job name, and what the endpoint does synchronously (shape check + enqueue).
    • Audit schema. Define the two tables from the architecture: imports
      (aggregate + current status) and import_events (append-only timeline with
      kind and a server-only detail). List the exact event kinds.
    • Storage. Decide where validated bytes live (object storage keyed by
      import id) and that the DB stores a pointer + metadata, not the blob.

    Done when: a BLUEPRINT document pins down the request/response shapes,
    the four ordered SSRF rules as a numbered list, the explicit sync-vs-job
    decision with rationale, the imports + import_events schema with named
    event kinds, and the storage choice. Every later task should be able to cite
    a line of this blueprint.

  3. Tasks — decompose into a DAG with minimal dependencies

    Break the blueprint into atomic, independently testable tasks and wire
    their dependencies into a DAG. The goal is a graph where anything that can
    be built in parallel is, and each edge is a real "needs the output of"
    relationship — not an accident of how you happened to think about it.

    Suggested decomposition (adapt to your stack):

    T1  URL validator        (scheme=https, no userinfo, non-empty host)
    T2  DNS + IP guard        (resolve host, block private/loopback/link-local)
    T3  Guarded fetch         (T1 + T2 + no auto-redirect, per-hop revalidate, limits)
    T4  Import endpoint       (shape check, create record, enqueue) — needs schema
    T5  Audit schema + writer (imports + import_events tables, event writer)
    T6  Import worker         (run T3, validate content, store, audit via T5)
    T7  Retry + idempotency   (safe re-run of T6 on failure)
    T8  Test suite            (proves every block from T1–T7 fails closed)
    

    Then draw the dependency edges and keep them minimal:

    flowchart LR
        T1[URL validator] --> T3[Guarded fetch]
        T2[DNS + IP guard] --> T3
        T5[Audit schema + writer] --> T4[Import endpoint]
        T5 --> T6[Import worker]
        T3 --> T6
        T4 --> T6
        T6 --> T7[Retry + idempotency]
        T7 --> T8[Test suite]
        T3 --> T8
    

    Notice the parallelism the graph exposes: T1, T2, and T5 have no
    dependencies
    and can be built at the same time. Keep the edges honest — if
    a task doesn't truly need another's output, don't add the edge, or you'll
    serialize work for no reason.

    Done when: you have a task list where each task is small enough to build
    and test on its own, a DAG (dare-dag.yaml or equivalent) that validates
    with no cycles and no broken references, and the graph makes the independent
    tasks (validator, IP guard, audit) visibly parallel. Each task should name
    its "done" check.

  4. Execute: SSRF validation — build the safe-fetch

    Implement the safe-fetch — the SSRF firewall from tasks T1–T3. This is
    the security core of the project; build it standalone and test it hard before
    wiring it into the worker. The companion SSRF lab walks the concepts;
    here you make it real in your repo.

    Build the pipeline in strict order — order is the security property:

    MAX_REDIRECTS = 2
    BLOCKED = [ "127.0.0.0/8","10.0.0.0/8","172.16.0.0/12","192.168.0.0/16",
                "169.254.0.0/16","0.0.0.0/8","::1/128","fc00::/7","fe80::/10" ]
    
    def safe_fetch(raw_url):
        url = raw_url
        for hop in 0..MAX_REDIRECTS:
            u = parse(url)                          # real parser, not regex
            if u.scheme != "https": reject          # kills http/file/gopher/ftp/data
            if u.userinfo present: reject           # https://ok@169.254.169.254/
            if u.host empty: reject
            ips = dns_resolve(u.host)               # may be several A/AAAA
            if ips empty: reject
            for ip in ips:                          # check EVERY resolved IP
                if numeric_in_cidr(ip, BLOCKED): reject   # incl. 169.254.169.254
            resp = http.get(u, follow_redirects=false, connect_to=ips, timeout=30s)
            if resp.is_redirect:
                url = resolve_relative(u, resp.header["Location"])
                continue                            # re-validate the NEW url
            return read_limited(resp)               # size + content-type limits
        reject("too many redirects")
    

    Non-negotiable details:

    • HTTPS-only, real parser, reject userinfo — an allowlist of one scheme
      beats a denylist you must keep complete; userinfo removes the
      trusted.com@internal ambiguity.
    • Block on the resolved IP, numerically — parse each IP to bytes and test
      CIDR containment. Immune to 0x7f.0.0.1, decimal 2130706433, and
      zero-padding. Cover IPv4-mapped IPv6 (::ffff:127.0.0.1). Reject if any
      resolved IP is internal. The 169.254.0.0/16 block (cloud metadata) is
      mandatory.
    • Drive redirects yourself, revalidate every hop — turn off auto-follow;
      a 302 → http://169.254.169.254/ must hit the full gauntlet again. Cap
      hops at ~2.
    • Limits — total timeout; max bytes checked on the header and by
      counting streamed bytes (Content-Length can lie); content-type allowlist;
      cap decompressed size to blunt decompression bombs.
    • Fail closed and quiet — every rejection returns a generic error; the
      reason is logged server-side only.

    Done when: safe_fetch exists as an isolated, unit-tested module that,
    with the network mocked, rejects: non-HTTPS schemes, a host resolving to
    169.254.169.254 / 10.0.0.5 / 127.0.0.1, a redirect to an internal IP,
    an oversized body (header and streamed), and a disallowed content-type — and
    allows a well-formed public HTTPS file. It never echoes upstream detail.

  5. Execute: async import — job, content validation, persistence & audit

    Wire the safe-fetch into the asynchronous import pipeline (tasks T4–T7).
    The endpoint enqueues; the worker does the risky work; every step is audited.
    The companion "Run and Diagnose Background Jobs" playbook is your
    reference for the worker, retries, and observability.

    Build it in two halves:

    1. The endpoint (synchronous, tiny).

    POST /imports (file_url):
        validate_shape(file_url)                    # present, string, length bound only
        imp = imports.create(user_id, file_url, status: "pending")
        audit(imp, kind: "queued")
        ImportJob.enqueue(imp.id)                    # returns before any fetch
        return 202, { id: imp.id, status: "pending" }
    

    2. The worker (asynchronous, guarded, audited).

    ImportJob(import_id):
        imp = imports.find(import_id)
        imp.update(status: "processing"); audit(imp, "fetch_started")
        try:
            bytes = safe_fetch(imp.file_url)         # milestone 4; may reject
        except Rejected as e:
            imp.update(status: "failed"); audit(imp, "fetch_rejected", detail: e.reason)
            return                                    # do NOT retry a policy rejection
        if not valid_content(bytes, expected_type):  # magic-byte sniff, CSV parse, row cap
            imp.update(status: "failed"); audit(imp, "content_rejected", detail: why)
            return
        key = storage.put(import_id, bytes)
        imp.update(status: "succeeded", storage_key: key,
                   content_type: sniffed, byte_size: len(bytes))
        audit(imp, "stored")
    

    Key decisions to get right:

    • Content validation is separate from transport limits. The safe-fetch
      bounds size/type at the wire; here you confirm the bytes are what you
      asked for
      — sniff magic bytes (don't trust the header), parse CSV and cap
      rows/columns, reject on parse failure.
    • Failure vs. retry. Distinguish a policy rejection (SSRF block, bad
      content — deterministic, will fail again → do not retry, mark failed)
      from a transient error (DNS blip, upstream 503, timeout → retry with
      backoff). Only retry the transient class.
    • Idempotent retries. A retried job must not double-store or double-audit.
      Key storage by import_id, make the "succeeded" transition a guarded
      state change, and record a retried event so the timeline stays truthful.
    • Audit every transition. queued → fetch_started → (fetch_rejected | content_rejected | stored | failed | retried). detail holds the
      server-only reason; the API still returns only a generic status.

    Done when: POST /imports enqueues and returns 202 without fetching;
    the worker runs the safe-fetch, validates content, stores the result, and
    writes an audit event for every transition; a policy rejection marks failed
    without retry while a transient error retries with backoff; and a retried job
    is idempotent (no duplicate storage or audit rows).

  6. Harden & verify — tests and submission criteria

    A security control you can't test will rot. Close the project by proving,
    with the network mocked, that every defense fails closed and the happy
    path works — then meet the submission bar.

    Cover, at minimum, these cases (mocking lets you simulate dangerous responses
    without real malicious infrastructure):

    test "rejects non-https scheme":
        expect_reject(safe_fetch("http://example.com/x"))
        expect_reject(safe_fetch("file:///etc/passwd"))
    
    test "rejects host resolving to a private / link-local IP":
        stub_dns("evil.test" => ["169.254.169.254"])      # cloud metadata
        expect_reject(safe_fetch("https://evil.test/"))
        stub_dns("lan.test"  => ["10.0.0.5"])
        expect_reject(safe_fetch("https://lan.test/"))
    
    test "rejects a redirect pointing at an internal IP":
        stub_http("https://ok.test/" => redirect_to("http://169.254.169.254/"))
        expect_reject(safe_fetch("https://ok.test/"))
    
    test "rejects an oversized file (header and streamed)":
        stub_http("https://big.test/" => body_of(500 MB))
        expect_reject(safe_fetch("https://big.test/"))
    
    test "the import job processes a valid file and audits it":
        stub_dns("cdn.test" => ["93.184.216.34"])          # public IP
        stub_http("https://cdn.test/a.csv" => csv_response(1 MB))
        imp = run_import("https://cdn.test/a.csv")
        expect(imp.status) == "succeeded"
        expect(audit_kinds(imp)) includes ["queued","fetch_started","stored"]
    
    test "a retry is idempotent (no duplicate storage or audit)":
        first  = run_import_failing_once("https://cdn.test/a.csv")  # transient error, then ok
        expect(imp.status) == "succeeded"
        expect(storage_objects(imp).count) == 1
        expect(audit_kinds(imp)) includes ["retried"]
    
    test "never leaks the upstream error to the caller":
        stub_http("https://ok.test/" => connection_refused())
        resp = post_imports("https://ok.test/")
        expect(resp.body) == generic_status()              # no timing/error oracle
    

    Also assert the cross-cutting rules: POST /imports returns before any fetch
    (the work is in a job), and the redirect hop cap is enforced.


    Submission criteria

    Submit when all of the following hold in your repo:

    1. Endpoint — POST /imports { file_url } validates shape only, creates a
      pending import, enqueues a background job, and returns 202 with an
      id; GET /imports/:id reports status with no upstream detail.
    2. Safe-fetch — enforces, in order: HTTPS-only + URL validation → DNS
      resolution blocking private/loopback/link-local (including
      169.254.169.254) → per-hop redirect revalidation (cap ~2) → size +
      timeout + content-type limits.
    3. Async import + audit — the worker runs the safe-fetch, validates
      content, stores the result, and writes an import_events row for every
      transition; policy rejections mark failed without retry, transient
      errors retry with backoff, and retries are idempotent.
    4. Tests (network mocked) proving each block: non-HTTPS, private/link-local
      IP (incl. metadata), malicious redirect, oversized file — plus a happy-path
      import that audits, an idempotent retry, and an "error is not leaked" test.

    Include a short note on DNS rebinding hardening (connect-by-IP with a
    pinned Host) as a future step, and confirm no upstream error, body, or
    timing is ever returned to the caller.