AI & Development
Step-by-step guide to building a production RAG (Retrieval-Augmented Generation) pipeline using Supabase, pgvector.
Fine-tuning a model on your data is expensive, slow, and produces a model that goes stale the moment your data updates. Retrieval-Augmented Generation (RAG) solves the same problem differently: at query time, retrieve the most relevant pieces of your data and include them in the model's context. The model answers using its own reasoning capabilities applied to your specific content. No training required, always up to date, and the retrieved sources are auditable. This guide builds a complete RAG pipeline using Supabase as the vector store and Claude as the generation model - the stack we use in production.
Every RAG system has the same four stages, though the specific implementation choices at each stage vary significantly:
The quality of a RAG system is determined almost entirely by the quality of the retrieval step. If the right content is not retrieved, the best LLM in the world cannot generate the right answer. Chunking strategy and retrieval configuration are where most RAG quality work happens.
Chunking - splitting documents into retrievable pieces - is where most RAG systems go wrong. The wrong chunk size results in either too little context per retrieved chunk (the answer is split across multiple chunks that are not all retrieved) or too much context per chunk (retrieval becomes imprecise because each chunk covers too many topics).
Practical chunking guidance:
function chunkDocument(text: string, chunkTokens = 300, overlapTokens = 75): string[] {
// Rough approximation: 1 token ≈ 4 characters
const chunkSize = chunkTokens * 4;
const overlapSize = overlapTokens * 4;
const chunks: string[] = [];
// Split on paragraph boundaries first
const paragraphs = text.split(/
+/);
let currentChunk = ', ';
for (const paragraph of paragraphs) {
if ((currentChunk + paragraph).length > chunkSize && currentChunk) {
chunks.push(currentChunk.trim());
// Start next chunk with overlap from end of current
currentChunk = currentChunk.slice(-overlapSize) + '
' + paragraph;
} else {
currentChunk += (currentChunk ? '
' : '') + paragraph;
}
}
if (currentChunk.trim()) chunks.push(currentChunk.trim());
return chunks;
}
import { createClient } from '@supabase/supabase-js';
import OpenAI from 'openai'; // Using OpenAI for embeddings (Claude doesn't yet have embedding models)
const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_KEY!);
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY! });
async function ingestDocument(doc: { title: string; url: string; content: string }) {
const chunks = chunkDocument(doc.content);
// Batch embed all chunks in one API call (up to 2,048 inputs)
const embeddingResponse = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: chunks,
});
const rows = chunks.map((chunk, i) => ({
content: chunk,
embedding: embeddingResponse.data[i].embedding,
metadata: {
source_title: doc.title,
source_url: doc.url,
chunk_index: i,
total_chunks: chunks.length,
},
}));
const { error } = await supabase.from('knowledge_chunks').insert(rows);
if (error) throw new Error(`Ingestion failed: ${error.message}`);
console.log(`Ingested ${chunks.length} chunks from "${doc.title}"`);
}
async function retrieveRelevantChunks(
question: string,
matchCount = 5,
matchThreshold = 0.7
): Promise<Chunk[]> {
// Embed the question
const questionEmbedding = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: question,
});
// Search for similar chunks
const { data, error } = await supabase.rpc('match_knowledge_chunks', {
query_embedding: questionEmbedding.data[0].embedding,
match_threshold: matchThreshold,
match_count: matchCount,
});
if (error) throw new Error(`Retrieval failed: ${error.message}`);
return data ?? [];
}
The Supabase RPC function (match_knowledge_chunks) is the pgvector similarity search wrapped in a database function - the same pattern described in the semantic search guide. The key parameter to tune is match_threshold: too low and you retrieve irrelevant chunks that confuse the model; too high and you miss relevant chunks that are phrased differently from the question.
async function answerQuestion(question: string): Promise<{ answer: string; sources: string[] }> {
const chunks = await retrieveRelevantChunks(question);
if (chunks.length === 0) {
return {
answer: "I don't have information about that in my knowledge base.",
sources: [],
};
}
// Assemble context
const context = chunks
.map((c, i) => `[Source ${i + 1}: ${c.metadata.source_title}]
${c.content}`)
.join('
---
');
const sources = [...new Set(chunks.map(c => c.metadata.source_url))];
// Generate answer
const anthropic = new Anthropic();
const response = await anthropic.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
system: `You are a helpful assistant that answers questions based only on the provided context.
If the context does not contain enough information to answer the question, say so clearly.
Always base your answer on the sources provided. Do not add information not present in the context.
When referencing specific information, mention which source it came from.`,
messages: [{
role: 'user',
content: `Context:
${context}
Question: ${question}`,
}],
});
return {
answer: response.content[0].text,
sources,
};
}
RAG built on Supabase is one of the most practical AI stacks available today: it requires no infrastructure beyond an existing Supabase project, it scales to millions of document chunks on a Pro plan, and the vector search and relational capabilities of Postgres let you combine semantic retrieval with standard SQL filters - filtering by date, by category, by access permissions - in a single query. That combination of flexibility and simplicity is why it is the default choice for teams building knowledge-base AI without dedicated ML infrastructure.