AI & Development
Evaluating LLM quality is harder than it looks. Learn the metrics that matter - from automated benchmarks to human evaluation
Evaluating a large language model is significantly harder than evaluating a traditional software system. A traditional API either returns the correct value or it does not. An LLM generates text, and text quality is multidimensional: it can be accurate but poorly formatted, or well-written but factually wrong, or helpful in general but wrong about a specific edge case. Building a reliable evaluation framework is as important as building the application itself - without it, you cannot know whether a model change improved things or made them worse.
LLM evaluation divides into two approaches: reference-based and reference-free. Reference-based evaluation compares the model's output to a known correct answer. This is straightforward for tasks with objective answers - question answering, code generation with tests, extraction tasks - and is the foundation of reliable automated evaluation. Reference-free evaluation assesses output quality without a comparison target, typically using an LLM as a judge or using human raters.
For most production applications, you need both. Reference-based evaluation catches factual errors and format violations reliably. Reference-free evaluation catches style issues, tone mismatches, and subjective quality problems that do not have an objective reference to compare against.
The right metric depends on the task. For information extraction tasks, precision (what fraction of extracted items are correct) and recall (what fraction of correct items were extracted) are the right metrics. For summarization, ROUGE scores compare n-gram overlap between the generated summary and a reference summary - they correlate reasonably well with human judgment on factual coverage but not on fluency or conciseness. For code generation, the only reliable metric is whether the code passes the tests.
Applying the wrong metric to a task is a common and costly mistake. ROUGE score on a creative writing task is meaningless. Exact match on a paraphrasing task is too strict. Choose metrics that match the actual quality dimensions that matter for your use case.
Using a capable LLM (typically GPT-4o or Claude) to evaluate the outputs of another LLM is now a standard technique for tasks where reference-based evaluation is not possible. The judge model receives the task, the model's output, and a rubric describing what good output looks like, and returns a score or a pass/fail judgment with reasoning.
LLM-as-judge is fast, scalable, and correlates reasonably well with human judgment on many tasks - studies consistently find agreement rates of 80% or higher between LLM judges and human raters on carefully designed rubrics. The limitations are real: LLM judges have position bias (preferring the first response in a comparison), verbosity bias (preferring longer responses), and self-enhancement bias (model A rates model A's outputs higher than model B's). Using the same model to judge its own outputs is especially unreliable.
The most important investment in LLM evaluation is curating a high-quality evaluation dataset. This is a set of inputs with known-correct or human-judged outputs that represents the full distribution of your production use cases, including edge cases and hard examples. Evaluation on a representative, challenging dataset tells you much more than evaluation on easy examples or synthetic data that was generated by the model being evaluated.
The evaluation dataset should be treated as a protected asset. If you use evaluation examples as training data (a practice called dataset contamination), the model may achieve high scores by memorization rather than generalization, and you will deploy a model that fails on real inputs not in the training set.
Automated evaluation metrics need to be calibrated against human judgment periodically. If your LLM-as-judge scores do not correlate with what your users actually think is a good response, the metric is misleading you. Periodic human evaluation - reviewing a sample of outputs and scoring them - keeps your automated metrics honest and helps you catch systematic biases in your evaluation setup.
For high-stakes applications, human evaluation is not optional. In medical, legal, or financial contexts, no automated metric should be the sole gate for production deployment. Human review of a representative sample before and after model changes is a necessary part of the deployment process.
LLM evaluation should run continuously, not just before deployment. Every change to the system prompt, the retrieval pipeline, the model version, or the post-processing logic is an opportunity for regression. Running your evaluation suite on every significant change and tracking metrics over time is the only way to catch quality degradation before it reaches users. This requires investment in evaluation infrastructure - a test runner, result storage, and metric visualization - but the return on that investment is a production system you can change with confidence.