AI Engineering · 90 min

Build a RAG Pipeline

Build a Retrieval-Augmented Generation pipeline from scratch: chunk a corpus, embed it, store vectors, run semantic search with an access filter, and generate grounded answers with source citations.

Get the starter project

Fork or clone it (Ruby/Rails, Python/FastAPI or TypeScript) and make the failing tests pass.

Problem

An LLM only knows what was in its training data. Ask it about your company's internal docs, a private knowledge base, or anything after its cutoff, and it guesses — confidently and wrong (a "hallucination"). **Retrieval-Augmented Generation (RAG)** fixes this by giving the model the right context *at question time*. Instead of relying on the model's memory, you retrieve the most relevant passages from your own corpus and paste them into the prompt. The model answers *from the evidence you provided* and cites its sources. In this lab you build the whole pipeline yourself — no framework doing the magic behind a single `.query()` call. You will chunk a corpus, embed it, store the vectors, run semantic search, enforce an access filter so restricted content never leaks, and finally generate a grounded, cited answer.

Objectives

Prerequisites

What is RAG, and why build it yourself?

RAG has four moving parts, and understanding each one is the point of this lab:

  1. Indexing (offline): turn your corpus into searchable vectors.
    Chunk → embed → store.
  2. Retrieval (per question): find the passages most similar in meaning to
    the question. Embed the question → search by cosine similarity → top-k.
  3. Augmentation: stuff those passages into the prompt as context.
  4. Generation: the LLM answers using that context and cites its sources.

The magic word is embeddings: a model maps text to a vector (a list of
numbers, e.g. 1536 of them) such that semantically similar text lands close
together
in that space. "How do I reset my password?" and "steps to recover
account access" are different words but nearby vectors. That's why semantic
search beats keyword search for questions.

A production RAG system is not just "call a library." The parts you have to get
right yourself — chunk size, caching, the metadata payload, the score threshold,
and above all the access filter — are exactly what this lab drills into.

You will build a small pipeline against a corpus of ~10–30 short documents and
end with a program that answers a question about that corpus, grounded in the
retrieved text and citing which documents it used.

Steps

  1. Prepare a corpus and chunk it with overlap

    Goal

    Gather a small corpus and split each document into overlapping chunks.

    Why chunk? Embedding models have a limited input window, and — more
    importantly — a vector for a whole 20-page document is a blurry average that
    matches nothing well. Smaller chunks give sharper, more retrievable vectors.

    Why overlap? If you cut on a hard boundary, a sentence explaining a
    concept can be split from the concept itself, and neither chunk answers the
    question. Overlap repeats the tail of one chunk at the head of the next so
    context survives the cut.

    Corpus

    Collect 10–30 short documents (Markdown or plain text is fine): FAQ entries,
    internal docs, blog posts. Mark at least a couple of them as restricted
    (premium: true) — you'll need them in Step 5.

    Chunking

    Use ~3000 characters (~800 tokens) per chunk with ~400 characters of
    overlap
    . Split on paragraph/sentence boundaries when you can, rather than
    mid-word.

    Document (9000 chars)
    |--------- chunk 0: chars 0..3000 ---------|
                            |--- overlap ---|
                      |--------- chunk 1: chars 2600..5600 ---------|
                                                    |--------- chunk 2 ---------|
    
    def chunk(text, size=3000, overlap=400):
        step = size - overlap          # 2600
        chunks = []
        start = 0
        while start < len(text):
            piece = text[start:start + size]
            chunks.append(piece)
            start += step
        return chunks
    

    Each chunk carries the identity of its parent document so you can cite it
    later. Model a chunk record like this:

    {
      "chunk_id": "doc-42:2",
      "doc_id": "doc-42",
      "title": "Password reset policy",
      "premium": false,
      "text": "…the chunk's ~3000 chars…"
    }
    

    Done when: running your chunker over the corpus produces a flat list of
    chunk records, each ≤ ~3000 chars, with overlap visible between consecutive
    chunks of the same document.

  2. Generate embeddings (with a cache)

    Goal

    Turn each chunk's text into a vector, and cache by content hash so you never
    pay to embed the same text twice.

    Embed

    Call your embeddings model on each chunk's text. With
    text-embedding-3-small you get a 1536-dimension vector per chunk. Note
    the dimension — the vector store collection must be created with exactly this
    number.

    vector = embed(chunk["text"])   # -> [0.013, -0.220, ..., 0.008]  (len 1536)
    

    Cache by content hash

    Embedding costs money and time. If a document didn't change, its chunks
    didn't change, so their vectors didn't change. Key a cache by a hash of the
    exact text:

    import hashlib
    
    def embed_cached(text, cache):
        key = hashlib.sha256(text.encode()).hexdigest()
        if key in cache:
            return cache[key]           # hit — no API call
        vec = embed(text)               # miss — call the model
        cache[key] = vec
        return vec
    

    The cache can be an in-memory dict for the lab, or a small table/file for
    real use. The point is the hash of the text is the key — change one
    character and you get a new key, so stale vectors can't linger.

    Batch

    Most embedding APIs accept a batch of inputs in one request. Send chunks in
    batches (e.g. 64 at a time) to cut latency and cost.

    Done when: every chunk record has a vector of the expected dimension,
    re-running the step is nearly free (all cache hits), and editing one document
    only re-embeds its chunks.

  3. Stand up a vector store and upsert with a payload

    Goal

    Run a vector store, create a collection with cosine distance, and upsert your
    vectors together with a metadata payload.

    Run it

    # Qdrant
    docker run -p 6333:6333 qdrant/qdrant
    # or pgvector
    docker run -p 5432:5432 -e POSTGRES_PASSWORD=pw pgvector/pgvector:pg16
    

    Create the collection

    The collection must match your embedding dimension and use cosine
    distance
    — the standard similarity metric for text embeddings.

    {
      "collection": "kb",
      "size": 1536,
      "distance": "Cosine"
    }
    

    Upsert points with a payload

    A "point" is a vector plus a payload — the metadata you'll filter on and
    cite from. Store exactly what retrieval and generation need: the source
    identity, the display title, the access flag, and the text itself.

    {
      "id": "doc-42:2",
      "vector": [0.013, -0.220, "…", 0.008],
      "payload": {
        "type": "doc",
        "doc_id": "doc-42",
        "title": "Password reset policy",
        "premium": false,
        "text": "…the chunk text, so you can cite and show it…"
      }
    }
    

    Index the payload fields you filter on (premium, type) if your store
    supports payload indexes — filtered search is much faster with them.

    Idempotent reindexing

    When a document changes and you reindex, delete its old points first,
    then insert the new ones. Otherwise stale chunks pile up and get retrieved
    forever.

    def reindex(doc):
        store.delete(collection="kb", filter={"doc_id": doc.id})   # remove old
        store.upsert(collection="kb", points=points_for(doc))      # insert new
    

    Done when: the collection reports the right vector count, and reindexing a
    document twice leaves the same number of points (no duplicates).

  4. Implement semantic search (embed, top-k, threshold)

    Goal

    Given a natural-language question, retrieve the most relevant chunks.

    The flow

    1. Embed the question with the same model you used for the chunks.
      This is non-negotiable — vectors from two different models don't share a
      space and their distances are meaningless.
    2. Search top-k by cosine similarity — ask the store for the k closest
      points (start with k = 5).
    3. Apply a score threshold to drop weak matches, so an off-topic question
      returns nothing instead of the least-irrelevant chunk.
    def search(question, k=5, min_score=0.35):
        qvec = embed(question)                       # same model as indexing!
        hits = store.search(
            collection="kb",
            vector=qvec,
            limit=k,
        )
        return [h for h in hits if h.score >= min_score]
    

    Example query and what comes back:

    Question: "How long until a reset link expires?"
    
    hits:
      0.71  doc-42:2  "Password reset policy"
      0.63  doc-42:1  "Password reset policy"
      0.41  doc-17:0  "Account security overview"
      0.22  doc-03:4  "Billing FAQ"          <- below threshold, dropped
    

    Tuning notes

    • k too low → you miss context the answer needed. k too high → you
      dilute the prompt with noise and pay for more tokens. k = 4–8 is a sane
      start.
    • Cosine scores run 0..1 for normalized embeddings; a good threshold is
      corpus-dependent, so measure a few real questions and pick a value that
      keeps good hits and drops junk.

    Done when: an on-topic question returns a short list of clearly relevant
    chunks (highest score first), and a nonsense question returns an empty list.

  5. Enforce an access filter in the query

    Goal

    Make sure a user without access can never retrieve restricted content — by
    filtering inside the vector-store query, not afterward.

    The filter

    Push a payload filter into the search itself. For a free user, restrict to
    premium = false:

    def search(question, user, k=5, min_score=0.35):
        qvec = embed(question)
        payload_filter = None if user.premium else {"premium": False}
        hits = store.search(
            collection="kb",
            vector=qvec,
            limit=k,
            filter=payload_filter,         # <- applied by the store, pre-ranking
        )
        return [h for h in hits if h.score >= min_score]
    

    In Qdrant this is a must condition on the payload; in pgvector it's a
    WHERE premium = false next to the ORDER BY embedding <=> $q — either way
    the store never returns the restricted points.

    Security note — why this is the whole ballgame

    Never retrieve everything and filter in application code afterward.
    Two ways that leaks:

    1. Restricted chunks eat your top-k slots, so even after you drop them the
      user gets a worse answer — a subtle denial of the content they should
      see.
    2. One missed branch — an error path, a debug log, a caching layer that
      stores the raw hits — and the restricted text is already in your
      process
      , one mistake away from the response.

    The filter belongs in the query so restricted vectors are never even
    scored. The database is your enforcement boundary. Treat the access flag as
    untrusted input from the user's session, not from the request body.

    Verify the leak is closed

    Ask a question whose best answer lives only in a premium: true document:

    • As a premium user → the restricted chunk is retrieved and cited.
    • As a free user → that chunk never appears; the answer falls back to
      public content or says it doesn't know.

    Done when: the exact same question returns different, correctly-scoped
    results for a premium vs. a free user, and the restricted text is provably
    absent from the free user's retrieved hits.

  6. Assemble the context prompt and generate a cited answer

    Goal

    Turn the retrieved chunks into a grounded, cited answer from the LLM — and
    ship it.

    Build the context prompt

    Concatenate the retrieved chunks, each tagged with its source title so the
    model can cite it. Give the model a firm instruction to answer only from
    the context and to admit when the context doesn't cover the question.

    def answer(question, hits):
        context = "\n\n".join(
            f"[Source: {h.payload['title']}]\n{h.payload['text']}"
            for h in hits
        )
        prompt = f"""Answer the question using ONLY the context below.
    If the context does not contain the answer, say you don't know.
    Cite the sources you used by their titles.
    
    Context:
    {context}
    
    Question: {question}
    """
        return llm(prompt)
    

    Return sources alongside the answer

    Don't rely only on the model to cite — you already know which documents fed
    the prompt, so return them as structured data too:

    {
      "answer": "A reset link expires after 30 minutes. [Password reset policy]",
      "sources": [
        { "doc_id": "doc-42", "title": "Password reset policy" }
      ]
    }
    

    Why "answer only from context"

    This instruction plus the retrieved evidence is what turns a hallucination
    machine into a grounded assistant. If retrieval returned nothing (empty after
    the threshold in Step 4), don't call the LLM with an empty context and let it
    improvise — return "I don't have information on that."

    Submission criterion

    Your pipeline answers a real question about your corpus, grounded in the
    retrieved chunks and citing the source titles — and it respects the access
    filter from Step 5 (a free user cannot get an answer built from restricted
    content).

    Deliver the repository, including:

    • The corpus (or a script that fetches it) with at least one premium: true document.
    • The full pipeline: chunk → embed (cached) → upsert → search → filter → generate.
    • A short README with the command to index the corpus and the command to ask a question.
    • A transcript (or test) showing: (a) a good question answered with citations,
      (b) the same premium-only question answered for a premium user but refused/
      redirected for a free user.