Build a RAG Pipeline
Build a Retrieval-Augmented Generation pipeline from scratch: chunk a corpus, embed it, store vectors, run semantic search with an access filter, and generate grounded answers with source citations.
Fork or clone it (Ruby/Rails, Python/FastAPI or TypeScript) and make the failing tests pass.
Problem
An LLM only knows what was in its training data. Ask it about your company's internal docs, a private knowledge base, or anything after its cutoff, and it guesses — confidently and wrong (a "hallucination"). **Retrieval-Augmented Generation (RAG)** fixes this by giving the model the right context *at question time*. Instead of relying on the model's memory, you retrieve the most relevant passages from your own corpus and paste them into the prompt. The model answers *from the evidence you provided* and cites its sources. In this lab you build the whole pipeline yourself — no framework doing the magic behind a single `.query()` call. You will chunk a corpus, embed it, store the vectors, run semantic search, enforce an access filter so restricted content never leaks, and finally generate a grounded, cited answer.
Objectives
- Split a text corpus into overlapping chunks that preserve context.
- Generate one embedding vector per chunk and cache by content hash to avoid re-embedding.
- Upsert vectors plus a metadata payload into a vector store using cosine distance.
- Implement semantic search: embed the question, retrieve top-k, apply a score threshold.
- Enforce an access/premium filter inside the query so restricted content never leaks.
- Assemble a context prompt from the retrieved passages and generate an answer with source citations.
- Make reindexing idempotent by deleting a document's old points before re-inserting.
Prerequisites
- An API key for an embeddings provider (e.g.
text-embedding-3-small, 1536 dims) or a local embedding model. - A vector store running locally — Qdrant or pgvector via Docker both work fine.
- An LLM you can call for the final generation step (any chat/completions model).
- The programming language of your choice; the lab is stack-agnostic and uses pseudocode.
- Basic comfort with HTTP/JSON APIs and running a container with
docker run.
What is RAG, and why build it yourself?
RAG has four moving parts, and understanding each one is the point of this lab:
-
Indexing (offline): turn your corpus into searchable vectors.
Chunk → embed → store. -
Retrieval (per question): find the passages most similar in meaning to
the question. Embed the question → search by cosine similarity → top-k. - Augmentation: stuff those passages into the prompt as context.
- Generation: the LLM answers using that context and cites its sources.
The magic word is embeddings: a model maps text to a vector (a list of
numbers, e.g. 1536 of them) such that semantically similar text lands close
together in that space. "How do I reset my password?" and "steps to recover
account access" are different words but nearby vectors. That's why semantic
search beats keyword search for questions.
A production RAG system is not just "call a library." The parts you have to get
right yourself — chunk size, caching, the metadata payload, the score threshold,
and above all the access filter — are exactly what this lab drills into.
You will build a small pipeline against a corpus of ~10–30 short documents and
end with a program that answers a question about that corpus, grounded in the
retrieved text and citing which documents it used.
Steps
-
Prepare a corpus and chunk it with overlap
Goal
Gather a small corpus and split each document into overlapping chunks.
Why chunk? Embedding models have a limited input window, and — more
importantly — a vector for a whole 20-page document is a blurry average that
matches nothing well. Smaller chunks give sharper, more retrievable vectors.Why overlap? If you cut on a hard boundary, a sentence explaining a
concept can be split from the concept itself, and neither chunk answers the
question. Overlap repeats the tail of one chunk at the head of the next so
context survives the cut.Corpus
Collect 10–30 short documents (Markdown or plain text is fine): FAQ entries,
internal docs, blog posts. Mark at least a couple of them as restricted
(premium: true) — you'll need them in Step 5.Chunking
Use ~3000 characters (~800 tokens) per chunk with ~400 characters of
overlap. Split on paragraph/sentence boundaries when you can, rather than
mid-word.Document (9000 chars) |--------- chunk 0: chars 0..3000 ---------| |--- overlap ---| |--------- chunk 1: chars 2600..5600 ---------| |--------- chunk 2 ---------|def chunk(text, size=3000, overlap=400): step = size - overlap # 2600 chunks = [] start = 0 while start < len(text): piece = text[start:start + size] chunks.append(piece) start += step return chunksEach chunk carries the identity of its parent document so you can cite it
later. Model a chunk record like this:{ "chunk_id": "doc-42:2", "doc_id": "doc-42", "title": "Password reset policy", "premium": false, "text": "…the chunk's ~3000 chars…" }Done when: running your chunker over the corpus produces a flat list of
chunk records, each ≤ ~3000 chars, with overlap visible between consecutive
chunks of the same document. -
Generate embeddings (with a cache)
Goal
Turn each chunk's text into a vector, and cache by content hash so you never
pay to embed the same text twice.Embed
Call your embeddings model on each chunk's
text. With
text-embedding-3-smallyou get a 1536-dimension vector per chunk. Note
the dimension — the vector store collection must be created with exactly this
number.vector = embed(chunk["text"]) # -> [0.013, -0.220, ..., 0.008] (len 1536)Cache by content hash
Embedding costs money and time. If a document didn't change, its chunks
didn't change, so their vectors didn't change. Key a cache by a hash of the
exact text:import hashlib def embed_cached(text, cache): key = hashlib.sha256(text.encode()).hexdigest() if key in cache: return cache[key] # hit — no API call vec = embed(text) # miss — call the model cache[key] = vec return vecThe cache can be an in-memory dict for the lab, or a small table/file for
real use. The point is the hash of the text is the key — change one
character and you get a new key, so stale vectors can't linger.Batch
Most embedding APIs accept a batch of inputs in one request. Send chunks in
batches (e.g. 64 at a time) to cut latency and cost.Done when: every chunk record has a
vectorof the expected dimension,
re-running the step is nearly free (all cache hits), and editing one document
only re-embeds its chunks. -
Stand up a vector store and upsert with a payload
Goal
Run a vector store, create a collection with cosine distance, and upsert your
vectors together with a metadata payload.Run it
# Qdrant docker run -p 6333:6333 qdrant/qdrant # or pgvector docker run -p 5432:5432 -e POSTGRES_PASSWORD=pw pgvector/pgvector:pg16Create the collection
The collection must match your embedding dimension and use cosine
distance — the standard similarity metric for text embeddings.{ "collection": "kb", "size": 1536, "distance": "Cosine" }Upsert points with a payload
A "point" is a vector plus a payload — the metadata you'll filter on and
cite from. Store exactly what retrieval and generation need: the source
identity, the display title, the access flag, and the text itself.{ "id": "doc-42:2", "vector": [0.013, -0.220, "…", 0.008], "payload": { "type": "doc", "doc_id": "doc-42", "title": "Password reset policy", "premium": false, "text": "…the chunk text, so you can cite and show it…" } }Index the payload fields you filter on (
premium,type) if your store
supports payload indexes — filtered search is much faster with them.Idempotent reindexing
When a document changes and you reindex, delete its old points first,
then insert the new ones. Otherwise stale chunks pile up and get retrieved
forever.def reindex(doc): store.delete(collection="kb", filter={"doc_id": doc.id}) # remove old store.upsert(collection="kb", points=points_for(doc)) # insert newDone when: the collection reports the right vector count, and reindexing a
document twice leaves the same number of points (no duplicates). -
Implement semantic search (embed, top-k, threshold)
Goal
Given a natural-language question, retrieve the most relevant chunks.
The flow
-
Embed the question with the same model you used for the chunks.
This is non-negotiable — vectors from two different models don't share a
space and their distances are meaningless. -
Search top-k by cosine similarity — ask the store for the
kclosest
points (start withk = 5). -
Apply a score threshold to drop weak matches, so an off-topic question
returns nothing instead of the least-irrelevant chunk.
def search(question, k=5, min_score=0.35): qvec = embed(question) # same model as indexing! hits = store.search( collection="kb", vector=qvec, limit=k, ) return [h for h in hits if h.score >= min_score]Example query and what comes back:
Question: "How long until a reset link expires?" hits: 0.71 doc-42:2 "Password reset policy" 0.63 doc-42:1 "Password reset policy" 0.41 doc-17:0 "Account security overview" 0.22 doc-03:4 "Billing FAQ" <- below threshold, droppedTuning notes
-
k too low → you miss context the answer needed. k too high → you
dilute the prompt with noise and pay for more tokens.k = 4–8is a sane
start. -
Cosine scores run 0..1 for normalized embeddings; a good threshold is
corpus-dependent, so measure a few real questions and pick a value that
keeps good hits and drops junk.
Done when: an on-topic question returns a short list of clearly relevant
chunks (highest score first), and a nonsense question returns an empty list. -
Embed the question with the same model you used for the chunks.
-
Enforce an access filter in the query
Goal
Make sure a user without access can never retrieve restricted content — by
filtering inside the vector-store query, not afterward.The filter
Push a payload filter into the search itself. For a free user, restrict to
premium = false:def search(question, user, k=5, min_score=0.35): qvec = embed(question) payload_filter = None if user.premium else {"premium": False} hits = store.search( collection="kb", vector=qvec, limit=k, filter=payload_filter, # <- applied by the store, pre-ranking ) return [h for h in hits if h.score >= min_score]In Qdrant this is a
mustcondition on the payload; in pgvector it's a
WHERE premium = falsenext to theORDER BY embedding <=> $q— either way
the store never returns the restricted points.Security note — why this is the whole ballgame
Never retrieve everything and filter in application code afterward.
Two ways that leaks:- Restricted chunks eat your top-k slots, so even after you drop them the
user gets a worse answer — a subtle denial of the content they should
see. - One missed branch — an error path, a debug log, a caching layer that
stores the raw hits — and the restricted text is already in your
process, one mistake away from the response.
The filter belongs in the query so restricted vectors are never even
scored. The database is your enforcement boundary. Treat the access flag as
untrusted input from the user's session, not from the request body.Verify the leak is closed
Ask a question whose best answer lives only in a
premium: truedocument:- As a premium user → the restricted chunk is retrieved and cited.
- As a free user → that chunk never appears; the answer falls back to
public content or says it doesn't know.
Done when: the exact same question returns different, correctly-scoped
results for a premium vs. a free user, and the restricted text is provably
absent from the free user's retrieved hits. - Restricted chunks eat your top-k slots, so even after you drop them the
-
Assemble the context prompt and generate a cited answer
Goal
Turn the retrieved chunks into a grounded, cited answer from the LLM — and
ship it.Build the context prompt
Concatenate the retrieved chunks, each tagged with its source title so the
model can cite it. Give the model a firm instruction to answer only from
the context and to admit when the context doesn't cover the question.def answer(question, hits): context = "\n\n".join( f"[Source: {h.payload['title']}]\n{h.payload['text']}" for h in hits ) prompt = f"""Answer the question using ONLY the context below. If the context does not contain the answer, say you don't know. Cite the sources you used by their titles. Context: {context} Question: {question} """ return llm(prompt)Return sources alongside the answer
Don't rely only on the model to cite — you already know which documents fed
the prompt, so return them as structured data too:{ "answer": "A reset link expires after 30 minutes. [Password reset policy]", "sources": [ { "doc_id": "doc-42", "title": "Password reset policy" } ] }Why "answer only from context"
This instruction plus the retrieved evidence is what turns a hallucination
machine into a grounded assistant. If retrieval returned nothing (empty after
the threshold in Step 4), don't call the LLM with an empty context and let it
improvise — return "I don't have information on that."Submission criterion
Your pipeline answers a real question about your corpus, grounded in the
retrieved chunks and citing the source titles — and it respects the access
filter from Step 5 (a free user cannot get an answer built from restricted
content).Deliver the repository, including:
- The corpus (or a script that fetches it) with at least one
premium: truedocument. - The full pipeline: chunk → embed (cached) → upsert → search → filter → generate.
- A short README with the command to index the corpus and the command to ask a question.
- A transcript (or test) showing: (a) a good question answered with citations,
(b) the same premium-only question answered for a premium user but refused/
redirected for a free user.
- The corpus (or a script that fetches it) with at least one