AI & Development

Building Semantic Search with pgvector and Embeddings (2026)

How to add semantic search to your app using pgvector and embedding models - from generating embeddings to writing similarity queries, with practical Supabase examples.

Keyword search fails at the edges. A user who searches "how to fix login not working" will not find an article titled "troubleshooting authentication errors" - even though the two phrases mean exactly the same thing. Semantic search closes this gap by comparing meaning rather than matching characters. The technical infrastructure behind it - embedding models and vector similarity search - has become accessible enough that you can add it to an existing Postgres database without a separate vector database service, using the pgvector extension. This guide walks through the complete implementation, from generating embeddings to serving results in production.

How semantic search works

An embedding model converts text into a dense vector - a list of floating-point numbers, typically 768 to 3072 values, that encodes the semantic meaning of the text. Texts with similar meanings produce vectors that are close together in this high-dimensional space. Semantic search works by:

  1. Generating an embedding vector for every document you want to search
  2. Storing those vectors in a database that supports vector similarity queries
  3. At search time, generating an embedding for the user's query
  4. Finding the stored document vectors closest to the query vector

The closeness metric is typically cosine similarity (the angle between vectors) or L2 distance (the Euclidean distance between vector endpoints). Cosine similarity is the standard for text embeddings - it is invariant to vector magnitude, which makes it more robust when comparing short queries to long documents.

Setting up pgvector in Supabase

Supabase Postgres has pgvector pre-installed. Enable it with:

-- In Supabase SQL editor or a migration file
create extension if not exists vector;

Then create a table with a vector column. The dimension must match your embedding model's output dimension:

create table documents (
  id bigserial primary key,
  title text not null,
  content text not null,
  embedding vector(1536), -- text-embedding-3-small is 1536 dimensions
  created_at timestamp with time zone default now()
);

-- Create an index for fast approximate nearest-neighbor search
create index on documents using ivfflat (embedding vector_cosine_ops)
  with (lists = 100); -- tune 'lists' based on your dataset size

The IVFFlat index enables approximate nearest-neighbor search, which is significantly faster than exact search for large datasets. For datasets under 100,000 rows, exact search (no index) is often fast enough and avoids the approximation error.

Generating and storing embeddings

Use an embedding model to convert your content to vectors before storing. For text, the most widely used models in 2026 are text-embedding-3-small (OpenAI, 1536 dims, $0.02/1M tokens) and text-embedding-ada-002 (OpenAI, 1536 dims, legacy). For a model-agnostic approach that avoids OpenAI dependency, nomic-embed-text (via Ollama or Nomic API, 768 dims) and mxbai-embed-large (1024 dims) are strong open-source alternatives.

import { createClient } from '@supabase/supabase-js';

const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_ANON_KEY!);
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY! });

async function embedAndStore(title: string, content: string) {
  // Generate embedding
  const embeddingResponse = await openai.embeddings.create({
    model: 'text-embedding-3-small',
    input: content.slice(0, 8000), // token limit safety
  });

  const embedding = embeddingResponse.data[0].embedding;

  // Store with embedding
  const { error } = await supabase.from('documents').insert({
    title,
    content,
    embedding,
  });

  if (error) throw error;
}

For bulk ingestion of large document sets, batch embedding requests - the embeddings API supports arrays of up to 2,048 input strings per request, which is dramatically more efficient than one-at-a-time calls.

Querying: the similarity search

Similarity search in pgvector uses the <=> (cosine distance), <-> (L2 distance), or <#> (negative inner product) operators. For text embeddings, use cosine distance:

-- SQL: find the 5 most similar documents to a query embedding
select
  id,
  title,
  content,
  1 - (embedding <=> '[0.023, -0.015...]'::vector) as similarity
from documents
order by embedding <=> '[0.023, -0.015...]'::vector
limit 5;

In practice, wrap this in a Supabase RPC function to avoid embedding the raw vector array in your application queries:

-- Supabase function definition
create or replace function search_documents(
  query_embedding vector(1536),
  match_threshold float default 0.7,
  match_count int default 5
)
returns table (id bigint, title text, content text, similarity float)
language sql stable
as $
  select
    id, title, content,
    1 - (embedding <=> query_embedding) as similarity
  from documents
  where 1 - (embedding <=> query_embedding) > match_threshold
  order by embedding <=> query_embedding
  limit match_count;
$;

The match_threshold parameter filters out results below a similarity score - typically 0.7 is a good starting point. Results below this threshold are usually genuinely irrelevant rather than weakly related.

Combining semantic and keyword search

Pure semantic search has a failure mode: it can find documents that are topically related but not actually relevant to a specific narrow query. "Where is the settings page?" is a semantic query that might return broad articles about customisation rather than the specific answer. Hybrid search - combining semantic similarity with keyword matching - handles both cases:

-- Hybrid search: semantic + keyword, reranked by combined score
select
  id, title, content,
  (
    0.7 * (1 - (embedding <=> query_embedding)) +  -- 70% semantic
    0.3 * ts_rank(to_tsvector('english', content), query_tsquery)  -- 30% keyword
  ) as combined_score
from documents, to_tsquery('english', $keyword_query) query_tsquery
where
  to_tsvector('english', content) @@ query_tsquery
  or (1 - (embedding <=> query_embedding)) > 0.6
order by combined_score desc
limit 10;

The weighting (70/30 semantic/keyword here) is a tunable parameter. Start at 70/30 and adjust based on your specific content and query patterns - content-heavy apps with long documents often benefit from more keyword weight, while conversational apps benefit from more semantic weight.

Performance at scale

pgvector with IVFFlat handles millions of rows with sub-100ms query times on a Supabase Pro instance. If you reach tens of millions of rows or need sub-10ms p99 latency, consider HNSW (hierarchical navigable small world) indexing - pgvector 0.5+ supports it and it is significantly faster than IVFFlat at scale:

-- HNSW index (better recall, faster queries at scale, higher memory use)
create index on documents using hnsw (embedding vector_cosine_ops)
  with (m = 16, ef_construction = 64);

For most applications serving under 5 million documents, IVFFlat or even no index (exact search) is sufficient. Optimise for correctness first, then performance if query times become an issue at your actual scale.

Semantic search is one of the highest-value AI features you can add to a content-heavy application - it fundamentally changes what users can find and how they interact with your content. The pgvector approach keeps everything in your existing Postgres database, which means you avoid a separate vector database service, maintain a single source of truth, and can combine vector search with all the relational queries you are already running.