AI & Development
How to build a production-ready AI backend on Firebase Cloud Functions - secure API key handling, streaming, per-user rate limiting.
Every mobile and web app that calls an LLM needs a secure backend proxy - the API key cannot live in client code, per-user rate limiting is not optional at scale, and you need cost visibility before the invoice arrives. Firebase Cloud Functions covers all of this without a dedicated server, without container management, and without a separate auth system. If you are already on Firebase for hosting, Firestore, or Auth, adding AI features through Cloud Functions is the path of least resistance. This guide covers the production architecture we use, including streaming, rate limiting, and cost tracking.
The short answer: your API key is a secret. App binaries can be decompiled. Embedding an API key in a mobile app or a frontend JS bundle exposes it to anyone who downloads the app - and API keys extracted from apps are routinely rotated across underground forums. Beyond the security issue, calling the API from the client means you cannot apply rate limiting, inject server-side context, log usage, or monitor costs per user. The backend proxy is not optional; the only question is which backend.
Firebase Cloud Functions wins for Firebase-first projects because authentication is seamless. Firebase Auth tokens are natively verified in Cloud Functions with one line of code, users are identified without any session management work, and the Functions deploy to the same project as your app's data.
A callable Cloud Function (invoked via the Firebase SDK rather than raw HTTP) handles auth verification automatically:
import { onCall, HttpsError } from 'firebase-functions/v2/https';
const anthropicKey = defineSecret('ANTHROPIC_API_KEY');
export const generateContent = onCall(
{ secrets: [anthropicKey], region: 'us-central1' },
async (request) => {
// Auth check - throws if not authenticated
if (!request.auth) {
throw new HttpsError('unauthenticated', 'Authentication required');
}
const { prompt, maxTokens = 1024 } = request.data;
const userId = request.auth.uid;
// Rate limit check (see below)
await checkRateLimit(userId);
const response = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-api-key': anthropicKey.value(),
'anthropic-version': '2023-06-01',
},
body: JSON.stringify({
model: 'claude-haiku-4-5-20251001',
max_tokens: maxTokens,
messages: [{ role: 'user', content: prompt }],
}),
});
if (!response.ok) {
throw new HttpsError('internal', `AI API error: ${response.status}`);
}
const data = await response.json();
return { content: data.content[0].text };
}
);
The defineSecret('ANTHROPIC_API_KEY') pattern stores the key in Firebase Secret Manager rather than environment variables. The key is never in source code, never in function configuration visible in the console, and rotated independently of the function deployment.
Without rate limiting, a single user can exhaust your API budget. The cleanest approach with Firebase is a daily token counter in Firestore, incremented transactionally with each request:
import { getFirestore, FieldValue } from 'firebase-admin/firestore';
const db = getFirestore();
const DAILY_TOKEN_LIMIT = 50000; // adjust per your cost budget
async function checkRateLimit(userId: string): Promise<void> {
const today = new Date().toISOString().slice(0, 10); // YYYY-MM-DD
const ref = db.doc(`usage/${userId}/daily/${today}`);
await db.runTransaction(async (tx) => {
const doc = await tx.get(ref);
const current = doc.exists ? doc.data()!.tokensUsed : 0;
if (current >= DAILY_TOKEN_LIMIT) {
throw new HttpsError('resource-exhausted', 'Daily AI limit reached. Resets at midnight UTC.');
}
tx.set(ref, { tokensUsed: FieldValue.increment(1) }, { merge: true });
});
}
async function recordUsage(userId: string, tokensUsed: number): Promise<void> {
const today = new Date().toISOString().slice(0, 10);
const ref = db.doc(`usage/${userId}/daily/${today}`);
await ref.set({ tokensUsed: FieldValue.increment(tokensUsed) }, { merge: true });
}
Call checkRateLimit before the API call and recordUsage after. The transactional check prevents race conditions where two simultaneous requests both see usage below the limit and both proceed past the check.
The callable function pattern does not support streaming - it returns a single response object. For streaming, use an HTTP function instead and stream the Anthropic SSE response directly to the client:
import { onRequest } from 'firebase-functions/v2/https';
export const streamContent = onRequest(
{ secrets: [anthropicKey], region: 'us-central1' },
async (req, res) => {
// Manual auth for HTTP functions
const token = req.headers.authorization?.replace('Bearer ''');
if (!token) { res.status(401).send('Unauthorized'); return; }
const decodedToken = await getAuth().verifyIdToken(token);
const userId = decodedToken.uid;
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
const upstream = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-api-key': anthropicKey.value(),
'anthropic-version': '2023-06-01',
},
body: JSON.stringify({
model: 'claude-sonnet-5',
max_tokens: 2048,
stream: true,
messages: [{ role: 'user', content: req.body.prompt }],
}),
});
// Pipe the SSE stream directly to the client
const reader = upstream.body!.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
res.write(decoder.decode(value));
}
res.end();
}
);
This pattern keeps the Cloud Function alive for the duration of the stream (typically 2 to 10 seconds for a typical response) and forwards tokens to the client as they arrive. Set the Cloud Function timeout to at least 60 seconds to accommodate long responses.
Firebase does not natively know your Anthropic costs, so you need to track them yourself. The usage documents you created for rate limiting are already there - add a cost estimate alongside the token count:
// After each API call, record usage + estimated cost
const inputTokens = data.usage.input_tokens;
const outputTokens = data.usage.output_tokens;
// Haiku pricing (per million tokens): input $0.80, output $4.00
const estimatedCostUsd = (inputTokens / 1_000_000 * 0.80) + (outputTokens / 1_000_000 * 4.00);
await db.doc(`usage/${userId}/daily/${today}`).set({
tokensUsed: FieldValue.increment(inputTokens + outputTokens),
estimatedCostUsd: FieldValue.increment(estimatedCostUsd),
}, { merge: true });
A simple Cloud Scheduler function that runs daily can aggregate these per-user costs into a top-level costs collection, giving you a daily cost dashboard in Firestore without any additional tooling.
Before deploying to production:
firebase functions:secrets:set ANTHROPIC_API_KEY{ cors: ['https://yourapp.com'] } - not true, which allows any originThe Firebase Cloud Functions architecture scales from zero to production with no infrastructure management and no operational overhead beyond the deployment. For solo developers and small teams shipping AI features on mobile or web, it is the right default backend - cheap at low volumes, scalable at high ones, and deeply integrated with every other Firebase service you are likely already using.