AI & Development
Standard product metrics miss what matters for AI features. Learn the AI-specific metrics - engagement, quality, coverage
Standard product analytics - daily active users, session length, retention - tells you whether users are using a feature. But for AI features, usage alone is not a useful signal. A user might interact with an AI chat feature every day for a week and then churn entirely when they realize it consistently gives them wrong answers. Or an AI feature might have low usage because its quality is excellent - users get their answer in one interaction and do not need to return. Standard metrics conflate quality and usage in ways that lead to wrong product decisions.
Coverage is the percentage of user requests that the AI handles completely, without the user needing to rephrase, retry, escalate to a human, or abandon the interaction. A coverage rate of 70% means 30% of users who use the AI feature do not get a satisfactory result from it. This is the most direct measure of whether the AI is doing its job.
Coverage is measured by tracking explicit signals (user escalated to human, user submitted a complaint, user rated the response negatively) and implicit signals (user sent the same or rephrased question again within the same session, suggesting the first answer was not satisfactory). Without coverage measurement, you cannot distinguish between "this AI feature is performing well" and "this AI feature is turning users away."
For AI features designed to help users accomplish specific goals - completing a form, finding information, resolving an issue - task completion rate is the primary success metric. Did the user accomplish the task they started with? This requires defining what task completion means for your specific use case and tracking whether it happens after an AI interaction.
Task completion is often measurable from downstream behavior: after a user interacts with an AI support agent, do they submit another support ticket about the same issue (suggesting the AI did not resolve it)? After using an AI search feature, do they reach the content they were looking for (measured by dwell time on the target page)? Connecting AI interactions to downstream outcomes requires event tracking that goes beyond the AI feature itself.
Explicit user feedback - thumbs up/down, star ratings, correction flows - provides direct quality signal. The ratio of positive to negative feedback is a leading indicator of quality issues. Analyzing the text of negative feedback (why did users rate a response poorly?) identifies specific failure patterns that inform prompt improvements or model changes.
Implicit quality signals supplement explicit feedback at higher volume. Short session duration after an AI interaction can indicate satisfaction (the user got what they needed quickly) or dissatisfaction (the user gave up). Repeat questions about the same topic often indicate the previous response did not answer it fully. Building a model that distinguishes between positive and negative implicit signals, calibrated against your explicit feedback data, extends quality measurement to all interactions rather than just the ones where users bother to give explicit feedback.
AI response latency directly affects user experience and completion rates. Users abandon interactions that feel too slow - mobile interactions especially. Measuring time-to-first-token (when the user sees the first character of the response) and total time-to-completion separately is important: streaming significantly improves the perceived latency even when wall-clock time is the same. Track latency by feature, by user type (connection speed matters), and by time of day (LLM provider latency varies under load). Alert on latency regressions promptly.
Unit economics for AI features require knowing the cost of each interaction, not just the total cost. Divide your AI API spend by the number of successful interactions (those that resulted in task completion or positive quality signals) to get the cost per successful interaction. This metric tells you whether improving quality (reducing failed interactions) also improves unit economics. An AI feature that handles 100% of requests poorly but cheaply is worse than one that handles 70% well at the same cost - the cost per successful interaction reveals this.
Tracking cost per successful interaction over time also reveals the impact of optimization work. A prompt improvement that increases coverage from 70% to 80% while keeping API cost constant reduces the cost per successful interaction by approximately 12.5% - a measurable, meaningful improvement that appears in the unit economics even if it does not change the total API spend.