AI & Development
Vector databases store and search high-dimensional embeddings for AI applications. Learn how they work, how they compare, and when you actually need one
A vector database is a database designed to store, index, and query high-dimensional numerical vectors - the embeddings that language models and image models use to represent semantic content. Traditional databases query by exact match or range: "find all rows where user_id = 42" or "find all products where price < 100." Vector databases query by similarity: "find the 10 vectors most similar to this query vector." This is the fundamental operation that makes semantic search, RAG, and recommendation systems work.
When you search a vector database, you provide a query vector - typically the embedding of a search query or document - and the database returns the k nearest neighbors: the stored vectors with the smallest distance to the query. Distance is measured using cosine similarity, dot product, or Euclidean distance, depending on the embedding model and use case.
Exact nearest-neighbor search requires comparing the query vector against every stored vector. This is feasible for small collections (tens of thousands of vectors) but becomes computationally expensive as collections grow into the millions or billions. Vector databases use approximate nearest neighbor (ANN) algorithms - HNSW, IVF, and others - to trade a small amount of accuracy for a massive speed improvement. Well-tuned ANN indexes return results in milliseconds even across hundreds of millions of vectors.
Hierarchical Navigable Small World (HNSW) is the indexing algorithm used by most production vector databases and most local vector search libraries. It builds a multi-layer graph of vectors where each node is connected to its nearest neighbors at that layer. Search navigates the graph from the top layer (coarse connections) down to the bottom layer (fine-grained connections), arriving at the approximate nearest neighbors efficiently.
HNSW is fast at query time but slower to build and more memory-intensive than some alternatives. For use cases where the dataset changes frequently, IVF (Inverted File Index) based approaches can offer better insert performance at the cost of somewhat lower recall. In practice, HNSW is the right choice for most applications unless dataset size or insert throughput forces a trade-off.
The major dedicated vector databases are Pinecone, Weaviate, Qdrant, and Chroma. Pgvector extends PostgreSQL with vector similarity search. Each has different trade-offs around performance at scale, ease of deployment, filtering capabilities, and cost.
Pgvector is the most practical choice for teams already running PostgreSQL who have modest scale requirements (up to a few million vectors). It avoids the operational overhead of a separate service and allows combining vector search with standard relational queries in a single query. Qdrant is a strong choice for self-hosted deployments where you need production-grade performance and full control. Pinecone is the managed option with the least operational overhead for teams that prefer to pay for convenience.
Real-world vector search almost always requires filtering by metadata alongside similarity. "Find the 10 most similar documents to this query, but only in the legal department's knowledge base, published after 2025." Vector databases vary significantly in how they implement this. Some filter before searching (pre-filtering), which requires loading large sets of filtered candidates before doing similarity search. Others filter after searching (post-filtering), which can miss relevant results if the filtered set is small. Hybrid approaches that integrate filtering into the ANN search itself - available in Qdrant and Weaviate - produce the best results for filtered queries.
For collections under 100,000 vectors, in-memory libraries like Faiss or even numpy-based nearest-neighbor search are often sufficient. The operational overhead of running a dedicated vector database service - monitoring, backups, version management - is not justified at small scale. Many early-stage AI applications are better served by starting with a simpler solution and migrating to a dedicated vector database only when performance or scale demands it.
Similarly, if your data fits in a PostgreSQL database and you are already running PostgreSQL, pgvector is almost always the right first choice. The performance characteristics are excellent for moderate scale, and the operational cost is zero if you are already managing Postgres.
The embedding model and the vector database must be aligned on embedding dimension. text-embedding-3-large produces 3072-dimensional vectors; text-embedding-3-small produces 1536-dimensional vectors. You cannot mix embeddings from different models in the same index, and changing embedding models requires re-embedding your entire dataset. This makes the choice of embedding model consequential: migrating later is expensive. Choose an embedding model with a track record of stability and evaluate it on your specific data before indexing at scale.