AI & Development

RAG Explained: Retrieval-Augmented Generation for Developers

Retrieval-Augmented Generation (RAG) lets LLMs answer questions from your own data without fine-tuning.

Retrieval-Augmented Generation, universally abbreviated as RAG, is a pattern for augmenting a language model's responses with information retrieved from an external data source at query time. Rather than baking knowledge into the model's weights through fine-tuning - an expensive, slow process - RAG retrieves relevant information dynamically and injects it into the prompt as context. The model answers based on what it retrieved, not just what it was trained on.

This pattern is now the dominant approach for building AI applications that need to answer questions from private or frequently updated data. Understanding how it works and where it breaks is essential for any developer building with LLMs.

The architecture

A RAG system has two phases: indexing and retrieval. During indexing, you take your data - documents, database records, support tickets, anything - chunk it into passages, generate embeddings for each chunk using an embedding model, and store those embeddings in a vector database. An embedding is a numerical vector that represents the semantic content of a passage: similar passages have vectors that are close in high-dimensional space.

During retrieval, when a user submits a query, you generate an embedding for the query, search the vector database for the most similar document chunks, retrieve those chunks, and inject them into the LLM prompt as context. The LLM then generates a response grounded in the retrieved content rather than generating from its training data alone.

Chunking strategy matters more than most tutorials suggest

The most underestimated variable in RAG quality is how you chunk your documents. If chunks are too small, each chunk lacks enough context for the LLM to give a coherent answer. If chunks are too large, you hit context window limits and include irrelevant information that dilutes the signal. The right chunk size depends on your data and your queries.

A common starting point is 512 tokens with a 50-token overlap between adjacent chunks. The overlap ensures that information at chunk boundaries is not lost when only one chunk is retrieved. Semantic chunking - splitting on paragraph or sentence boundaries rather than arbitrary token counts - generally outperforms fixed-size chunking because document structure carries meaning.

Embedding model choice

The embedding model determines how well semantic similarity search works. OpenAI's text-embedding-3-large and Cohere's embed-v3 are strong options for general-purpose English text. For multilingual use cases or domain-specific content, models fine-tuned on similar data outperform general-purpose models significantly.

Embedding quality is often the limiting factor in RAG recall - the fraction of relevant documents that actually get retrieved. If the embedding model does not understand your domain's terminology, queries will fail to retrieve relevant chunks even when those chunks contain exactly the information the user asked for. Evaluating retrieval recall on a curated set of question-answer pairs is the right way to measure whether your embedding choice is working.

Retrieval ranking

Vector similarity search returns the top-k most similar chunks, but similarity does not always correlate perfectly with relevance. Hybrid search - combining vector similarity with keyword search (BM25) - consistently outperforms pure vector search for RAG applications because it catches cases where exact keyword matches are more reliable than semantic similarity, such as product codes, names, or technical identifiers.

Re-ranking is a second retrieval step that takes the top-k candidates from the initial search and runs them through a cross-encoder model that scores each candidate against the query more carefully than the initial embedding comparison. Re-ranking adds latency but measurably improves precision - the fraction of retrieved chunks that are actually relevant.

When RAG fails

RAG has two failure modes: retrieval failure and generation failure. Retrieval failure happens when the right chunks are not retrieved - either because the query embedding does not match the document embedding well, because the relevant information was not indexed, or because the top-k was set too small. Generation failure happens when the right chunks are retrieved but the LLM does not use them correctly - ignoring them, contradicting them, or hallucinating details not present in the context.

The most common production issue is "lost in the middle": LLMs attend most strongly to the beginning and end of a long context, and chunks injected in the middle of a large context window are disproportionately likely to be ignored. Keeping retrieved context short and ordered by relevance mitigates this.

Evaluation is non-negotiable

Building a RAG system without evaluation is building blind. The minimum viable evaluation set is a list of questions with known correct answers drawn from your data. For each question, you check: did retrieval return the correct chunk? Did the LLM use that chunk to answer correctly? These two metrics - retrieval recall and answer accuracy - tell you where the system is failing and what to fix.

RAG evaluation frameworks like RAGAS provide automated metrics for faithfulness (does the answer match the context?), answer relevance, and context precision. Automated metrics are not a substitute for human evaluation on a representative sample, but they make continuous evaluation feasible as you iterate on chunking strategy, embedding models, and prompts.