AI & Development

Building a RAG Pipeline with Supabase: A Practical 2026 Guide

Step-by-step guide to building a production RAG (Retrieval-Augmented Generation) pipeline using Supabase, pgvector.

Fine-tuning a model on your data is expensive, slow, and produces a model that goes stale the moment your data updates. Retrieval-Augmented Generation (RAG) solves the same problem differently: at query time, retrieve the most relevant pieces of your data and include them in the model's context. The model answers using its own reasoning capabilities applied to your specific content. No training required, always up to date, and the retrieved sources are auditable. This guide builds a complete RAG pipeline using Supabase as the vector store and Claude as the generation model - the stack we use in production.

The RAG pipeline: four stages

Every RAG system has the same four stages, though the specific implementation choices at each stage vary significantly:

  1. Ingestion - Load your source documents, chunk them into retrievable pieces, generate embeddings for each chunk, store chunks and embeddings in the vector database
  2. Retrieval - When a user asks a question, generate an embedding for the question, find the most similar document chunks in the database
  3. Context assembly - Combine the retrieved chunks into a context window that can be passed to the LLM along with the question
  4. Generation - Call the LLM with the assembled context and the question; the model synthesises an answer from the provided content

The quality of a RAG system is determined almost entirely by the quality of the retrieval step. If the right content is not retrieved, the best LLM in the world cannot generate the right answer. Chunking strategy and retrieval configuration are where most RAG quality work happens.

Chunking strategy: the most underappreciated decision

Chunking - splitting documents into retrievable pieces - is where most RAG systems go wrong. The wrong chunk size results in either too little context per retrieved chunk (the answer is split across multiple chunks that are not all retrieved) or too much context per chunk (retrieval becomes imprecise because each chunk covers too many topics).

Practical chunking guidance:

  • Target 200 to 500 tokens per chunk - Small enough to be semantically coherent, large enough to contain a complete thought. A 300-token chunk is roughly 1 to 2 paragraphs.
  • Overlap chunks by 50 to 100 tokens - Overlap prevents relevant content that straddles a chunk boundary from being split between two chunks that might not both be retrieved.
  • Chunk by semantic structure, not character count - Split at paragraph boundaries, headers, and section breaks rather than at fixed character positions. A sentence that spans a chunk boundary is harder to use than a natural paragraph ending.
  • Preserve metadata in each chunk - Store the source document title, URL, section header, and position in the document alongside each chunk. Retrieval without provenance produces answers users cannot verify.
function chunkDocument(text: string, chunkTokens = 300, overlapTokens = 75): string[] {
  // Rough approximation: 1 token ≈ 4 characters
  const chunkSize = chunkTokens * 4;
  const overlapSize = overlapTokens * 4;
  const chunks: string[] = [];

  // Split on paragraph boundaries first
  const paragraphs = text.split(/

+/);
  let currentChunk = ', ';

  for (const paragraph of paragraphs) {
    if ((currentChunk + paragraph).length > chunkSize && currentChunk) {
      chunks.push(currentChunk.trim());
      // Start next chunk with overlap from end of current
      currentChunk = currentChunk.slice(-overlapSize) + '

' + paragraph;
    } else {
      currentChunk += (currentChunk ? '

' : '') + paragraph;
    }
  }

  if (currentChunk.trim()) chunks.push(currentChunk.trim());
  return chunks;
}

Ingestion pipeline: embedding and storing chunks

import { createClient } from '@supabase/supabase-js';

import OpenAI from 'openai'; // Using OpenAI for embeddings (Claude doesn't yet have embedding models)

const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_KEY!);
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY! });

async function ingestDocument(doc: { title: string; url: string; content: string }) {
  const chunks = chunkDocument(doc.content);

  // Batch embed all chunks in one API call (up to 2,048 inputs)
  const embeddingResponse = await openai.embeddings.create({
    model: 'text-embedding-3-small',
    input: chunks,
  });

  const rows = chunks.map((chunk, i) => ({
    content: chunk,
    embedding: embeddingResponse.data[i].embedding,
    metadata: {
      source_title: doc.title,
      source_url: doc.url,
      chunk_index: i,
      total_chunks: chunks.length,
    },
  }));

  const { error } = await supabase.from('knowledge_chunks').insert(rows);
  if (error) throw new Error(`Ingestion failed: ${error.message}`);

  console.log(`Ingested ${chunks.length} chunks from "${doc.title}"`);
}

Retrieval: finding the right chunks

async function retrieveRelevantChunks(
  question: string,
  matchCount = 5,
  matchThreshold = 0.7
): Promise<Chunk[]> {
  // Embed the question
  const questionEmbedding = await openai.embeddings.create({
    model: 'text-embedding-3-small',
    input: question,
  });

  // Search for similar chunks
  const { data, error } = await supabase.rpc('match_knowledge_chunks', {
    query_embedding: questionEmbedding.data[0].embedding,
    match_threshold: matchThreshold,
    match_count: matchCount,
  });

  if (error) throw new Error(`Retrieval failed: ${error.message}`);
  return data ?? [];
}

The Supabase RPC function (match_knowledge_chunks) is the pgvector similarity search wrapped in a database function - the same pattern described in the semantic search guide. The key parameter to tune is match_threshold: too low and you retrieve irrelevant chunks that confuse the model; too high and you miss relevant chunks that are phrased differently from the question.

Context assembly and generation

async function answerQuestion(question: string): Promise<{ answer: string; sources: string[] }> {
  const chunks = await retrieveRelevantChunks(question);

  if (chunks.length === 0) {
    return {
      answer: "I don't have information about that in my knowledge base.",
      sources: [],
    };
  }

  // Assemble context
  const context = chunks
    .map((c, i) => `[Source ${i + 1}: ${c.metadata.source_title}]
${c.content}`)
    .join('

---

');

  const sources = [...new Set(chunks.map(c => c.metadata.source_url))];

  // Generate answer
  const anthropic = new Anthropic();
  const response = await anthropic.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 1024,
    system: `You are a helpful assistant that answers questions based only on the provided context.
If the context does not contain enough information to answer the question, say so clearly.
Always base your answer on the sources provided. Do not add information not present in the context.
When referencing specific information, mention which source it came from.`,
    messages: [{
      role: 'user',
      content: `Context:
${context}

Question: ${question}`,
    }],
  });

  return {
    answer: response.content[0].text,
    sources,
  };
}

Common RAG failures and how to fix them

  • Retrieval misses the right chunk - The question uses different terminology than the source document. Fix: add query expansion (generate 2 to 3 alternative phrasings of the question and retrieve for each). Alternatively, use a re-ranker model to score retrieved chunks by relevance after initial retrieval.
  • Model ignores the context and uses parametric knowledge - The system prompt is not clear enough that answers must come from the provided context. Fix: add explicit grounding instructions - "Do not use any knowledge outside the provided context. If the answer is not in the context, say so."
  • Too many chunks overwhelm the context window - Retrieving 10 chunks of 500 tokens each is 5,000 tokens just for context. Fix: retrieve 5 chunks at most, rank them by similarity score, and truncate the lowest-scoring ones if the total context exceeds your target.
  • Stale embeddings after content updates - Documents that were updated still have old embeddings. Fix: store a content hash alongside the embedding and re-embed when the hash changes on re-ingestion.

RAG built on Supabase is one of the most practical AI stacks available today: it requires no infrastructure beyond an existing Supabase project, it scales to millions of document chunks on a Pro plan, and the vector search and relational capabilities of Postgres let you combine semantic retrieval with standard SQL filters - filtering by date, by category, by access permissions - in a single query. That combination of flexibility and simplicity is why it is the default choice for teams building knowledge-base AI without dedicated ML infrastructure.