Bootstrap a RAG Corpus
Operational guide to get your content into a live RAG system: configure the embedding and vector-store providers, index on publish via a callback, backfill the existing catalog, enforce access gating in the payload, and verify with a count plus a test semantic search.
PremiumThis is operations, not construction
Building a RAG pipeline from scratch — chunk → embed → vector store → search —
is a different job from running one in production. This playbook assumes the
pipeline already exists in your app. Your task here is to turn it on and
fill it up with real content, safely.
Four operational realities drive everything below:
-
Config comes from the environment. The embedding provider key and the
vector-store endpoint live in env/credentials, never in code. Get these wrong
and nothing indexes; leak them and you pay someone else's bill. -
Publishing is indexing. A publish event fires an
after_commitcallback
that enqueues an embedding job. You don't index by hand in normal operation —
you publish, and the content shows up in search a moment later. That "moment
later" only happens if the job worker is running (see the jobs playbook). -
The old catalog needs a backfill. Content that existed before RAG — or
that was published while the worker was down — never fired a live callback.
A one-time backfill enqueues indexing for every published item so the corpus
reflects your whole catalog, not just what you touched since launch. -
Access gating is a security boundary, not a display rule. Each indexed
point carries an access flag (e.g.premium) in its payload, and search
filters on it in the query. Restricted content must never surface in
semantic search for someone without access.
By the end you'll have a corpus that indexes automatically on publish, mirrors
your existing catalog, refuses to leak gated content, and that you can verify
with a chunk count and a real query.