Applied AI

Build a RAG-Powered Docs Assistant

A guided capstone: apply the DARE method end-to-end in your own repo to ship a docs assistant that ingests documents, embeds and indexes them, and answers questions with retrieval, access gating, and cited sources.

What you're building

This is a capstone project, not a starter you clone. You will build a
RAG-powered docs assistant from an empty repository, applying the DARE
method — Design, Blueprint, Tasks, Run/Execute — from the
first line to the last. The lab Build a RAG Pipeline taught you the moving
parts; the playbook Bootstrap a RAG Corpus taught you how to operate one.
Here you put both together into a real, verifiable system that you designed.

The finished assistant does one job well: a user asks a natural-language
question about a body of documents, and the assistant answers grounded in
those documents
, citing which ones it used — never inventing facts, and
never leaking content the user isn't allowed to see.

Under the hood that means six capabilities working as a pipeline:

documents ──▶ ingestion ──▶ chunking ──▶ embeddings ──▶ vector store
                                                             │
question ──▶ embed ──▶ retrieve (top-k, gated) ──▶ prompt ──▶ LLM ──▶ answer + sources

Why this project matters

RAG is the default architecture for putting a company's private knowledge
behind a chat interface: support bots, internal wikis, product docs, legal and
policy assistants. Anyone can wire a framework's .query() in an afternoon.
What separates a demo from something you'd put in front of users is the part
this project drills: a schema you can cite from, an access filter that
provably doesn't leak
, a retrieval step you can measure, and tests that
hold the line
. Those are engineering decisions, and DARE is how you make them
deliberately instead of by accident.

What you'll have at the end

How DARE guides the build

You don't start by writing pipeline code. You start by thinking, and DARE
gives that thinking a shape. Each phase produces an artifact the next phase
consumes:

The six milestones below map one-to-one onto this flow. Work them in order —
each one has a concrete "done when" criterion, and the whole point is that
you never write code you can't justify from the artifact above it.

Architecture

Arquitetura de referência

Mantenha o sistema como um conjunto de componentes pequenos com costuras
claras
, para que cada um possa ser construído e testado isoladamente. Seis
peças, dois fluxos.

Componentes

Fluxo de dados

INDEXAÇÃO (offline, ao publicar ou em backfill)
  documento ─▶ chunk(overlap) ─▶ embed(cache) ─▶ upsert{id,vector,payload}

RESPOSTA (por requisição)
  pergunta ─▶ embed ─▶ search(top-k, filter=acesso) ─▶ threshold
                                                          │
                    prompt(contexto + citações) ─▶ LLM ─▶ {resposta, fontes}

Diagrama do pipeline

flowchart TD
  subgraph Indexacao
    D[Documentos] --> C[Chunker com overlap]
    C --> E[Embedder cache por hash]
    E --> V[(Vector store cosseno)]
  end
  subgraph Resposta
    Q[Pergunta] --> QE[Embeddar pergunta]
    QE --> R[Retriever top-k]
    R -- filtro de acesso na query --> V
    V --> TH{score >= threshold?}
    TH -- nao --> EMPTY[Sem contexto relevante]
    TH -- sim --> P[Montar prompt contexto + fontes]
    P --> L[LLM caller]
    L --> A[Resposta + fontes citadas]
    EMPTY --> IDK[Nao sei responder]
  end

Decisões-chave e trade-offs

Milestones

  1. Design — frame the problem, users, and scope

    Apply the D of DARE

    Before any architecture, write a short DESIGN document (a page is
    plenty). This is where you decide what you're building and for whom —
    and, just as importantly, what you're leaving out. A RAG assistant with a
    fuzzy scope is impossible to test because "correct" was never defined.

    Answer these questions in writing

    • Users and access tiers. Who asks questions? Is there more than one
      tier of access (e.g. free vs. premium)? This decision drives the entire
      gating story later — name the tiers now.
    • Document sources. What goes into the corpus? Pick a concrete,
      bounded set for the capstone: 10–30 short documents (product docs, FAQ
      entries, policies). Mark at least two as restricted so you have
      something to gate.
    • Question types. What kinds of questions must it answer well
      ("how do I…", "what is the policy on…")? Write 5–8 real example
      questions — these become your evaluation set in Milestone 6.
    • Grounding contract. State the rule out loud: the assistant answers
      only from retrieved context and says "I don't know" when the corpus
      doesn't cover the question. No free-form world knowledge.
    • Non-goals. Write what this is not: not a general chatbot, not a
      web search, not a summarizer of documents it didn't retrieve, not
      multi-turn memory (unless you scope it in). Non-goals are what keep the
      project finishable.

    Deliverable

    A DESIGN.md in your repo capturing users, tiers, sources, question
    types, the grounding contract, and explicit non-goals.

    Done when: a reader who has never seen the project can state, from your
    DESIGN alone, who uses the assistant, what it will and won't answer, and
    what "a correct answer" means — including that restricted content is
    access-controlled.

  2. Blueprint — architecture, schema, endpoints, and the retriever contract

    Apply the B of DARE

    Turn the approved DESIGN into a BLUEPRINT: the concrete architecture
    the rest of the project builds against. This is where technical choices get
    locked in so tasks can proceed in parallel without stepping on each other.

    Choose your providers

    • Embedder + vector store. Pick one embedding model (note its
      dimension, e.g. 1536) and one vector store (Qdrant and pgvector both work
      locally via Docker). The collection must be created with that exact
      dimension and cosine distance — a mismatch is the #1 silent failure.
    • Chat model. A separate provider for generation, with its own key.

    Define the chunk / payload schema

    This schema is the contract every component shares — ingestion writes it,
    the retriever filters on it, the LLM caller cites from it. Nail it down now:

    {
      "id": "doc-42:2",
      "vector": [0.013, -0.220, "…"],
      "payload": {
        "doc_id": "doc-42",
        "title": "Password reset policy",
        "premium": false,
        "text": "…the chunk text, so you can cite and show it…"
      }
    }
    

    Every field earns its place: doc_id for idempotent reindex and citation,
    title for the cited source, premium (the access flag) for the query
    filter, text so you can build the prompt and show the evidence.

    Specify the endpoints and the retriever contract

    • POST /ask — request { question } (access comes from the session,
      never the body); response { answer, sources: [{doc_id, title}] }.
    • (Optional) index endpoint or task — how a document enters the corpus.
    • Retriever contract — freeze the signature so the API and LLM caller
      can be built against it:
    search(question, user, k=5, min_score=0.35) -> [Hit{ score, payload }]
      - embeds `question` with the SAME model as indexing
      - pushes an access filter into the store query when user is not premium
      - returns at most k hits with score >= min_score, ranked desc
    

    Deliverable

    A BLUEPRINT.md with the component list, chosen providers + dimension, the
    chunk/payload schema, the endpoint contracts, and the retriever signature.

    Done when: the schema, the POST /ask request/response shape, and the
    retriever signature are written down and stable enough that two people
    could build ingestion and answering separately and still fit together.

  3. Tasks — decompose into an atomic task DAG

    Apply the (t)A of DARE

    Break the blueprint into small, atomic tasks and arrange them into a
    DAG (directed acyclic graph) — a dependency graph that tells you what
    can be built now, what must wait, and what can run in parallel. A good task
    is one you can implement and test in a sitting, with a clear done signal.

    Suggested decomposition

    Each of these is one task; the arrows are dependencies:

    T1 chunker(text, size, overlap) -> [chunk]
    T2 embedder(text) -> vector          (cache by content hash)
    T3 vector-store client: create collection, upsert, delete, search
    T4 ingestion job: doc -> chunks -> embed -> upsert   (needs T1, T2, T3)
    T5 retriever.search(question, user, k, min_score)     (needs T2, T3)
    T6 prompt builder: hits -> context prompt
    T7 answer endpoint POST /ask: retrieve -> prompt -> LLM -> {answer, sources}  (needs T5, T6)
    

    The DAG

    flowchart LR
      T1[T1 chunker] --> T4[T4 ingestion job]
      T2[T2 embedder] --> T4
      T3[T3 vector store client] --> T4
      T2 --> T5[T5 retriever]
      T3 --> T5
      T5 --> T7[T7 answer endpoint]
      T6[T6 prompt builder] --> T7
    

    Order and parallelism

    • Rank 0 (parallel): T1, T2, T3, T6 — none depend on each other, so
      they can be built and unit-tested independently. T6 only shapes strings.
    • Rank 1: T4 (ingestion) and T5 (retriever) — each needs its rank-0
      dependencies but not the other, so they too can proceed in parallel.
    • Rank 2: T7 (the answer endpoint) — the join point; it needs the
      retriever and the prompt builder.

    This ordering is why Milestones 4 and 5 split the way they do: everything
    the ingestion+index wave needs (T1–T4) is independent of the
    retrieval+answer wave (T5–T7) except through the vector store and the
    embedder, which are shared, frozen contracts.

    Deliverable

    A TASKS.md (or a DAG file) listing each atomic task, its dependencies,
    and its done criterion.

    Done when: every capability in the blueprint maps to at least one task,
    every task lists its dependencies, and the graph has no cycles — you can
    read off a valid build order and see which tasks are parallelizable.

  4. Execute — ingestion + index

    Run the first execution wave

    Implement tasks T1–T4: get documents into the vector store as clean,
    idempotent, gated points. Build each piece against the frozen schema, test
    it in isolation, then wire them into one ingestion path.

    Chunk with overlap (T1)

    Split each document into overlapping chunks that preserve context across
    the cut. Attach the parent identity to every chunk so you can cite it later.

    step = size - overlap            # e.g. 3000 - 400 = 2600
    for start in 0, step, 2*step, …:
        emit text[start : start + size]  with {doc_id, title, premium}
    

    Embed with a cache (T2)

    Embed each chunk's text with your chosen model. Key a cache on
    sha256(text) so re-running is nearly free and editing one document only
    re-embeds its chunks. Batch requests (e.g. 64 at a time) to cut cost and
    latency. Assert every vector has the expected dimension.

    Index into the vector store (T3, T4)

    Create the collection with the exact embedding dimension and cosine
    distance. Upsert points carrying the full payload (doc_id, title,
    premium, text). Make reindex idempotent — delete a document's
    existing points before inserting the new ones:

    reindex(doc):
        store.delete(filter={doc_id: doc.id})   # remove old
        store.upsert(points_for(doc))           # insert fresh
    

    Stamp the access flag correctly at index time: sub-content inherits its
    parent's gating
    (a lesson in a premium course is premium), so resolve the
    effective premium and never leave it null.

    Backfill the corpus

    Run a one-time pass over every document to populate the store, reusing the
    same ingestion path. Because reindex is idempotent, a backfill that dies
    halfway can be re-run with no duplicates.

    Done when: running the backfill produces a collection whose point
    count
    matches your expectation (≈ documents × average chunks), reindexing
    a document twice leaves that count unchanged (no duplicates), and every
    point carries a correct premium flag. Confirm the count — "the job didn't
    error" is not proof; an empty collection also doesn't error.

  5. Execute — retrieval + answer (with gating)

    Run the second execution wave

    Implement tasks T5–T7: given a question, retrieve the right gated chunks,
    build a grounded prompt, call the LLM, and return an answer with sources.

    Similarity search, top-k, threshold (T5)

    Embed the question with the same model used for indexing — vectors from
    two different models don't share a space. Ask the store for the top-k
    nearest by cosine, then drop anything below a score threshold so an
    off-topic question returns nothing rather than the least-irrelevant chunk.

    Gating: filter premium BEFORE top-k (the critical part)

    Push the access filter into the store query, so restricted vectors are
    never even scored:

    search(question, user, k=5, min_score=0.35):
        qvec = embed(question)                     # same model as indexing
        filter = None if user.premium else {premium: False}
        hits = store.search(vector=qvec, limit=k, filter=filter)   # pre-ranking
        return [h for h in hits if h.score >= min_score]
    

    Why before top-k, not after. If you retrieve everything and drop
    restricted hits in application code, two things break: restricted chunks
    steal top-k slots (the free user gets a worse answer — a subtle denial of
    content they should see), and the restricted text is already in your
    process
    , one missed branch away from the response. The database is the
    enforcement boundary. Derive user.premium from the session, never
    from the request body.

    Build the context prompt and generate (T6, T7)

    Concatenate the retrieved chunks, each tagged with its source title, and
    instruct the model to answer only from the context and to admit when the
    context doesn't cover the question. If retrieval returned nothing after the
    threshold, don't call the LLM with empty context — return "I don't have
    information on that."

    Return sources as structured data, not just inline text — you already know
    which documents fed the prompt:

    {
      "answer": "A reset link expires after 30 minutes. [Password reset policy]",
      "sources": [ { "doc_id": "doc-42", "title": "Password reset policy" } ]
    }
    

    Done when: POST /ask answers an on-topic question grounded in
    retrieved chunks and returns the sources it used; an off-topic question
    yields "I don't know" instead of a hallucination; and the same question
    whose best answer lives only in a premium document is answered for a premium
    user but not for a free user — with the restricted text provably absent from
    the free user's retrieved hits.

  6. Harden & verify — tests, limits, and the submission criterion

    Run to green, then prove it

    The pipeline works on a happy path — now make it trustworthy. This
    milestone is the R/verify closing of DARE: tests that hold the line,
    attention to limits, and a written definition of "done".

    Write the tests that matter

    Three properties, each a test that would fail loudly if a future change
    broke it:

    • Retrieval relevance. For each example question from your DESIGN
      (Milestone 1), the expected document appears in the top hits. A nonsense
      question returns an empty result (the threshold works).
    • Gating doesn't leak. The same question whose answer lives only in
      a premium document: a premium user retrieves and cites it; a free user's
      hits provably do not contain the restricted text. This is the test that
      matters most — it encodes a security boundary, not a nicety.
    • The answer cites its source. A successful answer returns a non-empty
      sources array whose doc_ids actually fed the prompt, and the answer
      text references them.

    Mind the limits

    • Latency. POST /ask does an embed + a vector search + an LLM call.
      Measure it; the LLM call dominates. Don't embed inline on ingestion inside
      a web request — that path is asynchronous.
    • Empty and error paths. No hits after threshold → "I don't know",
      never a call to the LLM with empty context. A provider timeout →
      a clean error, not a leaked stack trace.
    • Cost. Embedding is a real, metered cost. The content-hash cache and
      batching (from Milestone 4) keep re-indexing cheap; confirm re-publishing
      unchanged content produces cache hits, not new charges.
    • Dimension mismatch. Re-assert the collection dimension equals the
      embedder's — it's the #1 silent failure and worth a startup health check.

    Submission criterion — what "done" means

    Deliver the repository, containing:

    • The DARE artifacts: DESIGN.md, BLUEPRINT.md, TASKS.md (+ DAG).
    • The corpus (or a script that fetches it) with at least one restricted
      (premium: true) document.
    • The full pipeline: chunk → embed (cached) → upsert → search → gate →
      generate, wired to a working POST /ask.
    • A short README with the command to index/backfill the corpus and the
      command (or request) to ask a question.
    • The passing test suite, plus a transcript showing: (a) a good question
      answered with citations, (b) the same premium-only question answered for
      a premium user but refused/redirected for a free user.

    Done when: the three tests pass, the point count and a live query
    verify the corpus is populated and retrievable, the gate demonstrably
    changes results between a free and a premium user, and a reviewer can index
    the corpus and ask a question using only your README.