Build a RAG-Powered Docs Assistant
A guided capstone: apply the DARE method end-to-end in your own repo to ship a docs assistant that ingests documents, embeds and indexes them, and answers questions with retrieval, access gating, and cited sources.
What you're building
This is a capstone project, not a starter you clone. You will build a
RAG-powered docs assistant from an empty repository, applying the DARE
method — Design, Blueprint, Tasks, Run/Execute — from the
first line to the last. The lab Build a RAG Pipeline taught you the moving
parts; the playbook Bootstrap a RAG Corpus taught you how to operate one.
Here you put both together into a real, verifiable system that you designed.
The finished assistant does one job well: a user asks a natural-language
question about a body of documents, and the assistant answers grounded in
those documents, citing which ones it used — never inventing facts, and
never leaking content the user isn't allowed to see.
Under the hood that means six capabilities working as a pipeline:
documents ──▶ ingestion ──▶ chunking ──▶ embeddings ──▶ vector store
│
question ──▶ embed ──▶ retrieve (top-k, gated) ──▶ prompt ──▶ LLM ──▶ answer + sources
Why this project matters
RAG is the default architecture for putting a company's private knowledge
behind a chat interface: support bots, internal wikis, product docs, legal and
policy assistants. Anyone can wire a framework's .query() in an afternoon.
What separates a demo from something you'd put in front of users is the part
this project drills: a schema you can cite from, an access filter that
provably doesn't leak, a retrieval step you can measure, and tests that
hold the line. Those are engineering decisions, and DARE is how you make them
deliberately instead of by accident.
What you'll have at the end
- A repository you own, structured by the DARE artifacts (a design doc, an
architecture blueprint, a task DAG, and an execution log). - A working
POST /askendpoint (or CLI equivalent) that returns a grounded
answer plus a structured list of sources. - An ingestion path that chunks documents with overlap, embeds them, and
indexes them into a vector store — idempotently. - Access gating enforced inside the retrieval query, verified by a test
that proves a free user can never retrieve premium content. - A test suite covering relevance, gating, and citation, plus a written
submission criterion that defines "done".
How DARE guides the build
You don't start by writing pipeline code. You start by thinking, and DARE
gives that thinking a shape. Each phase produces an artifact the next phase
consumes:
-
Design — frame the problem, the users, and the scope. What is this
assistant, and just as importantly, what is it not? -
Blueprint — turn the design into an architecture: the components, the
chunk/payload schema, the endpoints, and the retriever contract. -
Tasks — decompose the blueprint into small, atomic tasks and order them
into a dependency graph (a DAG) so you always know what to build next. -
Run / Execute — implement the tasks in two waves (ingestion+index, then
retrieval+answer), then harden and verify until the tests prove it works.
The six milestones below map one-to-one onto this flow. Work them in order —
each one has a concrete "done when" criterion, and the whole point is that
you never write code you can't justify from the artifact above it.
Architecture
Arquitetura de referência
Mantenha o sistema como um conjunto de componentes pequenos com costuras
claras, para que cada um possa ser construído e testado isoladamente. Seis
peças, dois fluxos.
Componentes
-
Ingestão — lê os documentos-fonte (Markdown/texto/HTML), normaliza-os e
quebra cada um em chunks com overlap. É dona das fronteiras de chunk e da
identidade do documento pai. -
Embedder — transforma o texto de um chunk em um vetor usando um modelo de
embedding. Cacheia por hash do conteúdo para que texto inalterado nunca seja
re-embeddado. A dimensão que ele emite (ex.: 1536) é um contrato rígido com o
vector store. -
Vector store — guarda pontos
{id, vector, payload}e roda a busca por
similaridade de cosseno com um filtro de payload. Esta é a sua fronteira de
enforcement para o controle de acesso. -
Retriever — embedda uma pergunta com o mesmo modelo, aplica o filtro de
acesso, pede ao store os top-k, e descarta os matches abaixo de um threshold
de score. Retorna hits rankeados, não uma resposta. -
LLM caller — monta um prompt de contexto a partir dos chunks recuperados,
instrui o modelo a responder apenas a partir do contexto, e chama um modelo
de chat (separado). -
API / UI — expõe
POST /ask, orquestra retriever → LLM caller, e retorna
{ answer, sources }. Deriva o acesso do usuário a partir da sessão.
Fluxo de dados
INDEXAÇÃO (offline, ao publicar ou em backfill)
documento ─▶ chunk(overlap) ─▶ embed(cache) ─▶ upsert{id,vector,payload}
RESPOSTA (por requisição)
pergunta ─▶ embed ─▶ search(top-k, filter=acesso) ─▶ threshold
│
prompt(contexto + citações) ─▶ LLM ─▶ {resposta, fontes}
Diagrama do pipeline
flowchart TD
subgraph Indexacao
D[Documentos] --> C[Chunker com overlap]
C --> E[Embedder cache por hash]
E --> V[(Vector store cosseno)]
end
subgraph Resposta
Q[Pergunta] --> QE[Embeddar pergunta]
QE --> R[Retriever top-k]
R -- filtro de acesso na query --> V
V --> TH{score >= threshold?}
TH -- nao --> EMPTY[Sem contexto relevante]
TH -- sim --> P[Montar prompt contexto + fontes]
P --> L[LLM caller]
L --> A[Resposta + fontes citadas]
EMPTY --> IDK[Nao sei responder]
end
Decisões-chave e trade-offs
-
Tamanho de chunk e overlap. Chunks menores geram vetores mais nítidos mas
perdem o contexto ao redor; o overlap recompra esse contexto ao custo de um
pouco de duplicação. Comece com ~3000 chars / ~400 de overlap e ajuste contra
queries reais. -
top-k. Baixo demais e você deixa o prompt sem o trecho que continha a
resposta; alto demais e você o dilui com ruído e paga por tokens.k = 4–8é
uma faixa inicial sensata; um threshold de score te protege de retornar o
chunk "menos irrelevante" em perguntas fora do tema. -
Gating no payload, filtrado na query. A flag de acesso vive no payload de
cada ponto e o filtro é empurrado para dentro da busca, para que vetores
restritos nem sejam pontuados. Filtrar depois do retrieval é um vazamento
(chunks restritos roubam vagas de top-k, e o texto cru já está no seu
processo). Derive a flag da sessão, nunca do corpo da requisição. -
Dois provedores, mantidos separados. O modelo de embedding e o modelo de
chat são serviços diferentes com chaves diferentes, para você trocar cada um
de forma independente. -
Indexação idempotente. Reindexar deleta os pontos antigos de um documento
antes de inserir os novos, para que callbacks de publicação e backfills possam
rodar quantas vezes for e caírem no mesmo estado limpo. -
O contrato do retriever. Congele a assinatura do retriever cedo —
search(pergunta, usuario, k, min_score) -> [Hit{score, payload}]— para que
a API e o LLM caller possam ser construídos contra ela antes de ela estar
totalmente implementada.
Milestones
-
Design — frame the problem, users, and scope
Apply the D of DARE
Before any architecture, write a short DESIGN document (a page is
plenty). This is where you decide what you're building and for whom —
and, just as importantly, what you're leaving out. A RAG assistant with a
fuzzy scope is impossible to test because "correct" was never defined.Answer these questions in writing
-
Users and access tiers. Who asks questions? Is there more than one
tier of access (e.g. free vs. premium)? This decision drives the entire
gating story later — name the tiers now. -
Document sources. What goes into the corpus? Pick a concrete,
bounded set for the capstone: 10–30 short documents (product docs, FAQ
entries, policies). Mark at least two as restricted so you have
something to gate. -
Question types. What kinds of questions must it answer well
("how do I…", "what is the policy on…")? Write 5–8 real example
questions — these become your evaluation set in Milestone 6. -
Grounding contract. State the rule out loud: the assistant answers
only from retrieved context and says "I don't know" when the corpus
doesn't cover the question. No free-form world knowledge. -
Non-goals. Write what this is not: not a general chatbot, not a
web search, not a summarizer of documents it didn't retrieve, not
multi-turn memory (unless you scope it in). Non-goals are what keep the
project finishable.
Deliverable
A
DESIGN.mdin your repo capturing users, tiers, sources, question
types, the grounding contract, and explicit non-goals.Done when: a reader who has never seen the project can state, from your
DESIGN alone, who uses the assistant, what it will and won't answer, and
what "a correct answer" means — including that restricted content is
access-controlled. -
Users and access tiers. Who asks questions? Is there more than one
-
Blueprint — architecture, schema, endpoints, and the retriever contract
Apply the B of DARE
Turn the approved DESIGN into a BLUEPRINT: the concrete architecture
the rest of the project builds against. This is where technical choices get
locked in so tasks can proceed in parallel without stepping on each other.Choose your providers
-
Embedder + vector store. Pick one embedding model (note its
dimension, e.g. 1536) and one vector store (Qdrant and pgvector both work
locally via Docker). The collection must be created with that exact
dimension and cosine distance — a mismatch is the #1 silent failure. - Chat model. A separate provider for generation, with its own key.
Define the chunk / payload schema
This schema is the contract every component shares — ingestion writes it,
the retriever filters on it, the LLM caller cites from it. Nail it down now:{ "id": "doc-42:2", "vector": [0.013, -0.220, "…"], "payload": { "doc_id": "doc-42", "title": "Password reset policy", "premium": false, "text": "…the chunk text, so you can cite and show it…" } }Every field earns its place:
doc_idfor idempotent reindex and citation,
titlefor the cited source,premium(the access flag) for the query
filter,textso you can build the prompt and show the evidence.Specify the endpoints and the retriever contract
-
POST /ask— request{ question }(access comes from the session,
never the body); response{ answer, sources: [{doc_id, title}] }. - (Optional) index endpoint or task — how a document enters the corpus.
-
Retriever contract — freeze the signature so the API and LLM caller
can be built against it:
search(question, user, k=5, min_score=0.35) -> [Hit{ score, payload }] - embeds `question` with the SAME model as indexing - pushes an access filter into the store query when user is not premium - returns at most k hits with score >= min_score, ranked descDeliverable
A
BLUEPRINT.mdwith the component list, chosen providers + dimension, the
chunk/payload schema, the endpoint contracts, and the retriever signature.Done when: the schema, the
POST /askrequest/response shape, and the
retriever signature are written down and stable enough that two people
could build ingestion and answering separately and still fit together. -
Embedder + vector store. Pick one embedding model (note its
-
Tasks — decompose into an atomic task DAG
Apply the (t)A of DARE
Break the blueprint into small, atomic tasks and arrange them into a
DAG (directed acyclic graph) — a dependency graph that tells you what
can be built now, what must wait, and what can run in parallel. A good task
is one you can implement and test in a sitting, with a clear done signal.Suggested decomposition
Each of these is one task; the arrows are dependencies:
T1 chunker(text, size, overlap) -> [chunk] T2 embedder(text) -> vector (cache by content hash) T3 vector-store client: create collection, upsert, delete, search T4 ingestion job: doc -> chunks -> embed -> upsert (needs T1, T2, T3) T5 retriever.search(question, user, k, min_score) (needs T2, T3) T6 prompt builder: hits -> context prompt T7 answer endpoint POST /ask: retrieve -> prompt -> LLM -> {answer, sources} (needs T5, T6)The DAG
flowchart LR T1[T1 chunker] --> T4[T4 ingestion job] T2[T2 embedder] --> T4 T3[T3 vector store client] --> T4 T2 --> T5[T5 retriever] T3 --> T5 T5 --> T7[T7 answer endpoint] T6[T6 prompt builder] --> T7
Order and parallelism
-
Rank 0 (parallel): T1, T2, T3, T6 — none depend on each other, so
they can be built and unit-tested independently. T6 only shapes strings. -
Rank 1: T4 (ingestion) and T5 (retriever) — each needs its rank-0
dependencies but not the other, so they too can proceed in parallel. -
Rank 2: T7 (the answer endpoint) — the join point; it needs the
retriever and the prompt builder.
This ordering is why Milestones 4 and 5 split the way they do: everything
the ingestion+index wave needs (T1–T4) is independent of the
retrieval+answer wave (T5–T7) except through the vector store and the
embedder, which are shared, frozen contracts.Deliverable
A
TASKS.md(or a DAG file) listing each atomic task, its dependencies,
and its done criterion.Done when: every capability in the blueprint maps to at least one task,
every task lists its dependencies, and the graph has no cycles — you can
read off a valid build order and see which tasks are parallelizable. -
Rank 0 (parallel): T1, T2, T3, T6 — none depend on each other, so
-
Execute — ingestion + index
Run the first execution wave
Implement tasks T1–T4: get documents into the vector store as clean,
idempotent, gated points. Build each piece against the frozen schema, test
it in isolation, then wire them into one ingestion path.Chunk with overlap (T1)
Split each document into overlapping chunks that preserve context across
the cut. Attach the parent identity to every chunk so you can cite it later.step = size - overlap # e.g. 3000 - 400 = 2600 for start in 0, step, 2*step, …: emit text[start : start + size] with {doc_id, title, premium}Embed with a cache (T2)
Embed each chunk's text with your chosen model. Key a cache on
sha256(text)so re-running is nearly free and editing one document only
re-embeds its chunks. Batch requests (e.g. 64 at a time) to cut cost and
latency. Assert every vector has the expected dimension.Index into the vector store (T3, T4)
Create the collection with the exact embedding dimension and cosine
distance. Upsert points carrying the full payload (doc_id,title,
premium,text). Make reindex idempotent — delete a document's
existing points before inserting the new ones:reindex(doc): store.delete(filter={doc_id: doc.id}) # remove old store.upsert(points_for(doc)) # insert freshStamp the access flag correctly at index time: sub-content inherits its
parent's gating (a lesson in a premium course is premium), so resolve the
effectivepremiumand never leave it null.Backfill the corpus
Run a one-time pass over every document to populate the store, reusing the
same ingestion path. Because reindex is idempotent, a backfill that dies
halfway can be re-run with no duplicates.Done when: running the backfill produces a collection whose point
count matches your expectation (≈ documents × average chunks), reindexing
a document twice leaves that count unchanged (no duplicates), and every
point carries a correctpremiumflag. Confirm the count — "the job didn't
error" is not proof; an empty collection also doesn't error. -
Execute — retrieval + answer (with gating)
Run the second execution wave
Implement tasks T5–T7: given a question, retrieve the right gated chunks,
build a grounded prompt, call the LLM, and return an answer with sources.Similarity search, top-k, threshold (T5)
Embed the question with the same model used for indexing — vectors from
two different models don't share a space. Ask the store for the top-k
nearest by cosine, then drop anything below a score threshold so an
off-topic question returns nothing rather than the least-irrelevant chunk.Gating: filter premium BEFORE top-k (the critical part)
Push the access filter into the store query, so restricted vectors are
never even scored:search(question, user, k=5, min_score=0.35): qvec = embed(question) # same model as indexing filter = None if user.premium else {premium: False} hits = store.search(vector=qvec, limit=k, filter=filter) # pre-ranking return [h for h in hits if h.score >= min_score]Why before top-k, not after. If you retrieve everything and drop
restricted hits in application code, two things break: restricted chunks
steal top-k slots (the free user gets a worse answer — a subtle denial of
content they should see), and the restricted text is already in your
process, one missed branch away from the response. The database is the
enforcement boundary. Deriveuser.premiumfrom the session, never
from the request body.Build the context prompt and generate (T6, T7)
Concatenate the retrieved chunks, each tagged with its source title, and
instruct the model to answer only from the context and to admit when the
context doesn't cover the question. If retrieval returned nothing after the
threshold, don't call the LLM with empty context — return "I don't have
information on that."Return sources as structured data, not just inline text — you already know
which documents fed the prompt:{ "answer": "A reset link expires after 30 minutes. [Password reset policy]", "sources": [ { "doc_id": "doc-42", "title": "Password reset policy" } ] }Done when:
POST /askanswers an on-topic question grounded in
retrieved chunks and returns the sources it used; an off-topic question
yields "I don't know" instead of a hallucination; and the same question
whose best answer lives only in a premium document is answered for a premium
user but not for a free user — with the restricted text provably absent from
the free user's retrieved hits. -
Harden & verify — tests, limits, and the submission criterion
Run to green, then prove it
The pipeline works on a happy path — now make it trustworthy. This
milestone is the R/verify closing of DARE: tests that hold the line,
attention to limits, and a written definition of "done".Write the tests that matter
Three properties, each a test that would fail loudly if a future change
broke it:-
Retrieval relevance. For each example question from your DESIGN
(Milestone 1), the expected document appears in the top hits. A nonsense
question returns an empty result (the threshold works). -
Gating doesn't leak. The same question whose answer lives only in
a premium document: a premium user retrieves and cites it; a free user's
hits provably do not contain the restricted text. This is the test that
matters most — it encodes a security boundary, not a nicety. -
The answer cites its source. A successful answer returns a non-empty
sourcesarray whosedoc_ids actually fed the prompt, and the answer
text references them.
Mind the limits
-
Latency.
POST /askdoes an embed + a vector search + an LLM call.
Measure it; the LLM call dominates. Don't embed inline on ingestion inside
a web request — that path is asynchronous. -
Empty and error paths. No hits after threshold → "I don't know",
never a call to the LLM with empty context. A provider timeout →
a clean error, not a leaked stack trace. -
Cost. Embedding is a real, metered cost. The content-hash cache and
batching (from Milestone 4) keep re-indexing cheap; confirm re-publishing
unchanged content produces cache hits, not new charges. -
Dimension mismatch. Re-assert the collection dimension equals the
embedder's — it's the #1 silent failure and worth a startup health check.
Submission criterion — what "done" means
Deliver the repository, containing:
- The DARE artifacts:
DESIGN.md,BLUEPRINT.md,TASKS.md(+ DAG). - The corpus (or a script that fetches it) with at least one restricted
(premium: true) document. - The full pipeline: chunk → embed (cached) → upsert → search → gate →
generate, wired to a workingPOST /ask. - A short README with the command to index/backfill the corpus and the
command (or request) to ask a question. - The passing test suite, plus a transcript showing: (a) a good question
answered with citations, (b) the same premium-only question answered for
a premium user but refused/redirected for a free user.
Done when: the three tests pass, the point count and a live query
verify the corpus is populated and retrievable, the gate demonstrably
changes results between a free and a premium user, and a reviewer can index
the corpus and ask a question using only your README. -
Retrieval relevance. For each example question from your DESIGN