AI & Development

Rolling Out AI Features Safely with Feature Flags (2026 Guide)

AI features have unique rollout risks: unpredictable outputs, cost spikes, latency variance.

Shipping an AI feature is not like shipping a CRUD feature. A bug in a standard feature either works or does not - binary, reproducible, fixable with a patch. An AI feature can work correctly for 99% of inputs and produce something surprising or harmful for the remaining 1%. It can spike in cost when a prompt pattern appears that you did not anticipate. It can change behaviour silently when a model provider pushes an update. These failure modes are not theoretical - every team that has shipped AI features in production has encountered at least one of them. Feature flags are how you ship confidently despite those unknowns.

What makes AI feature rollouts different

Standard feature flags give you: percentage rollout, user segment targeting, A/B test assignment, and instant kill switch. For AI features, you need all of those plus a few additional dimensions:

  • Model version pinning - LLM providers update models without notice. The behaviour that worked yesterday may not work exactly the same today. Being able to pin to a specific model version via a flag means provider updates do not hit production silently.
  • Cost circuit breakers - AI inference costs can spike unexpectedly. A flag that disables the feature if daily cost exceeds a threshold prevents surprises on your billing statement.
  • Fallback configuration - When the AI feature is disabled (by a kill switch or a circuit breaker), what does the user see? Flags should configure the fallback experience, not just turn the feature off.
  • Prompt versioning - Prompt changes are code changes that affect output. Being able to roll back a prompt change without a full deployment is the same benefit as rolling back code, and flags are the right mechanism.

Firebase Remote Config for AI feature flags

Firebase Remote Config is well-suited for AI feature flags because it supports typed parameter groups, conditional targeting, and A/B test assignment - and is already available if you use Firebase. A typical AI feature flag configuration in Remote Config:

// Remote Config parameter group: "ai_assistant"
{
  "ai_assistant_enabled": true,           // master kill switch
  "ai_model": "claude-haiku-4-5-20251001",  // pinned model version
  "ai_rollout_percentage": 25,            // % of users who get the feature
  "ai_max_tokens": 1024,                  // cost control
  "ai_fallback_mode": "static",           // what non-AI users see
  "ai_system_prompt_version": "v3",       // prompt version to use
  "ai_cost_circuit_breaker_usd": 50.0     // daily cost ceiling
}

Fetch and apply in your app:

import { getRemoteConfig, fetchAndActivate, getValue } from 'firebase/remote-config';

const remoteConfig = getRemoteConfig(app);
remoteConfig.defaultConfig = {
  ai_assistant_enabled: false,   // safe default: off
  ai_model: 'claude-haiku-4-5-20251001',
  ai_rollout_percentage: 0,
  ai_max_tokens: 512,
  ai_fallback_mode: 'static',
};

await fetchAndActivate(remoteConfig);

const aiEnabled = getValue(remoteConfig'ai_assistant_enabled').asBoolean();
const rolloutPct = getValue(remoteConfig'ai_rollout_percentage').asNumber();

// Stable user assignment (same user always gets same bucket)
const userBucket = hashUserId(currentUser.uid) % 100;
const isInRollout = aiEnabled && userBucket < rolloutPct;

The stable hashing ensures the same user always gets the same experience - a user who had the AI feature enabled yesterday does not lose it today just because they hit a different bucket on the next fetch.

Implementing a cost circuit breaker

A circuit breaker pattern for AI costs checks daily spend before each request and disables the feature if a threshold is exceeded - protecting you from a traffic spike or a prompt pattern that generates unexpectedly long responses:

// In your Cloud Function, before calling the AI API
async function checkCostCircuitBreaker(): Promise<void> {
  const today = new Date().toISOString().slice(0, 10);
  const costDoc = await db.doc(`costs/daily/${today}`).get();
  const spent = costDoc.data()?.totalUsd ?? 0;

  const ceiling = parseFloat(
    getValue(remoteConfig'ai_cost_circuit_breaker_usd').asString()
  );

  if (spent >= ceiling) {
    // Log the circuit breaker event for alerting
    await db.doc(`alerts/circuit_breaker`).set({
      triggeredAt: new Date().toISOString(),
      spent,
      ceiling,
    }, { merge: true });

    throw new Error(`AI feature temporarily disabled: daily cost ceiling reached (${ceiling})`);
  }
}

Set up a Firestore trigger on the alerts/circuit_breaker document to send you an email or Slack message when this fires. The circuit breaker buys you time to investigate the cause without an open-ended cost run.

Prompt versioning strategy

Prompt changes should be versioned and flagged just like code changes. Store prompt versions in Firestore rather than hardcoding them in your function:

// Fetch the active prompt version (set by Remote Config flag)
async function getSystemPrompt(): Promise<string> {
  const version = getValue(remoteConfig'ai_system_prompt_version').asString();
  const doc = await db.doc(`prompts/system/versions/${version}`).get();

  if (!doc.exists) {
    // Fall back to v1 if version not found
    const fallback = await db.doc('prompts/system/versions/v1').get();
    return fallback.data()?.content ?? defaultPrompt;
  }

  return doc.data()!.content;
}

This pattern means rolling back a bad prompt change is a Remote Config update (seconds), not a code deployment (minutes). It also means you can A/B test prompt variants as easily as you A/B test UI changes - assign users to prompt version groups via Remote Config conditions and measure the output quality difference.

Gradual rollout sequence

A responsible AI feature rollout follows a progression:

  1. Internal users only (0.1%) - Your team, opted-in beta testers. Verify the feature works end-to-end and that output quality meets expectations.
  2. 5% rollout - First real users. Monitor error rates, latency p95, and API costs closely for 48 hours.
  3. 25% rollout - Validate that the cost and latency numbers scale linearly from 5%. If a 5× increase in traffic produces a 10× increase in costs, there is a bug in your token usage that needs fixing before going wider.
  4. 50% → 100% - Two-stage expansion to full rollout, with 24-hour monitoring windows at each stage.

The kill switch - setting ai_assistant_enabled: false in Remote Config - should be accessible to any team member in under 30 seconds. This is not a corner case; it is a capability you should expect to use at least once per feature, and the ease of using it determines how quickly you can respond to a production issue.

What to monitor during rollout

  • Error rate - API call failures, validation failures, circuit breaker triggers
  • Latency p95 - LLM calls introduce significant tail latency; watch p95, not just mean
  • Cost per user per day - Aggregate from your usage tracking; verify it matches budget estimates
  • Feature engagement rate - Are users in the rollout cohort actually using the feature?
  • Downstream impact - Is the feature affecting the metrics it was designed to improve (retention, session length, conversion)?

Feature flags are not a temporary scaffolding to remove after launch. The best AI features stay behind flags permanently - because the ability to swap model versions, adjust token limits, and kill the feature in response to a provider incident is too valuable to give up just because the feature shipped successfully.