AI & Development
The TypeScript patterns that make LLM-powered applications maintainable at scale - typed clients, middleware chains, retry logic, provider abstractions.
Most TypeScript AI applications start the same way: a direct API call, a response.content[0].text, and a quick JSON.parse. That works until the first production incident - a provider outage, a model version change that subtly alters output format, a cost spike from a prompt pattern nobody anticipated. The teams that recover quickly from those incidents have one thing in common: they built with enough abstraction that changing one thing (the model, the provider, the retry logic, the cost tracking) does not require touching every feature that uses AI. This guide covers the TypeScript patterns that make that possible.
If your application calls the Anthropic SDK directly in 15 different places, switching providers or adding fallback logic requires 15 changes. A provider abstraction wraps the SDK call behind a common interface that all features use:
// Types
interface LLMRequest {
model?: string;
system?: string;
messages: Array<{ role: 'user' | 'assistant'; content: string }>;
maxTokens?: number;
temperature?: number;
}
interface LLMResponse {
content: string;
inputTokens: number;
outputTokens: number;
model: string;
provider: string;
}
// Provider interface
interface LLMProvider {
complete(request: LLMRequest): Promise<LLMResponse>;
stream(request: LLMRequest): AsyncIterable<string>;
}
// Anthropic implementation
class AnthropicProvider implements LLMProvider {
private client: Anthropic;
private defaultModel: string;
constructor(apiKey: string, defaultModel = 'claude-sonnet-5') {
this.client = new Anthropic({ apiKey });
this.defaultModel = defaultModel;
}
async complete(request: LLMRequest): Promise<LLMResponse> {
const response = await this.client.messages.create({
model: request.model ?? this.defaultModel,
max_tokens: request.maxTokens ?? 1024,
system: request.system,
messages: request.messages,
});
return {
content: response.content[0].type === 'text' ? response.content[0].text : '',
inputTokens: response.usage.input_tokens,
outputTokens: response.usage.output_tokens,
model: response.model,
provider: 'anthropic',
};
}
async *stream(request: LLMRequest): AsyncIterable<string> {
const stream = await this.client.messages.stream({
model: request.model ?? this.defaultModel,
max_tokens: request.maxTokens ?? 1024,
system: request.system,
messages: request.messages,
});
for await (const chunk of stream) {
if (chunk.type === 'content_block_delta' && chunk.delta.type === 'text_delta') {
yield chunk.delta.text;
}
}
}
}
Features call provider.complete() or provider.stream() - never the Anthropic client directly. When you add OpenAI as a fallback or switch to a different model family, you add a new provider implementation. Existing feature code is untouched.
Retry logic, cost tracking, logging, and rate limiting all need to run around every AI call - but they should not be duplicated in every feature. A middleware chain wraps the provider and applies these behaviours automatically:
type LLMMiddleware = (
request: LLMRequest,
next: (req: LLMRequest) => Promise<LLMResponse>
) => Promise<LLMResponse>;
function withRetry(maxAttempts = 3, baseDelayMs = 500): LLMMiddleware {
return async (request, next) => {
let lastError: Error | null = null;
for (let attempt = 0; attempt < maxAttempts; attempt++) {
try {
return await next(request);
} catch (err) {
lastError = err as Error;
if (attempt < maxAttempts - 1) {
const delay = baseDelayMs * Math.pow(2, attempt) + Math.random() * 100;
await new Promise(r => setTimeout(r, delay));
}
}
}
throw lastError;
};
}
function withCostTracking(trackFn: (tokens: { input: number; output: number; model: string }) => void): LLMMiddleware {
return async (request, next) => {
const response = await next(request);
trackFn({ input: response.inputTokens, output: response.outputTokens, model: response.model });
return response;
};
}
function withLogging(logger: { info: (...args: any[]) => void }): LLMMiddleware {
return async (request, next) => {
const start = Date.now();
try {
const response = await next(request);
logger.info('LLM call succeeded', {
model: response.model,
inputTokens: response.inputTokens,
outputTokens: response.outputTokens,
durationMs: Date.now() - start,
});
return response;
} catch (err) {
logger.info('LLM call failed', { error: (err as Error).message, durationMs: Date.now() - start });
throw err;
}
};
}
// Compose middleware
function buildLLMClient(provider: LLMProvider, middlewares: LLMMiddleware[]): LLMProvider {
const composed = (request: LLMRequest): Promise<LLMResponse> => {
const chain = middlewares.reduceRight(
(next, middleware) => (req: LLMRequest) => middleware(req, next),
(req: LLMRequest) => provider.complete(req)
);
return chain(request);
};
return { complete: composed, stream: provider.stream.bind(provider) };
}
Usage:
const llm = buildLLMClient(
new AnthropicProvider(process.env.ANTHROPIC_API_KEY!),
[
withRetry(3, 500),
withCostTracking((tokens) => costTracker.record(tokens)),
withLogging(logger),
]
);
// All features use this single client
const result = await llm.complete({ messages: [{ role: 'user', content: prompt }] });
Prompts are code. Treating them as template strings scattered across the codebase makes them hard to test, version, and review. Typed prompt builders centralise prompt construction and make the variables explicit:
interface SummarisationPromptVars {
content: string;
maxWords: number;
tone: 'formal' | 'casual' | 'technical';
audience: string;
}
const summarisationPrompt = {
system: (vars: Pick<SummarisationPromptVars'tone' | 'audience'>) =>
`You are a ${vars.tone} summariser writing for ${vars.audience}.
Return only the summary. No preamble. No "Here is a summary:" header.`,
user: (vars: SummarisationPromptVars) =>
`Summarise the following content in no more than ${vars.maxWords} words:
${vars.content}`,
};
// Usage - TypeScript enforces all required variables
const response = await llm.complete({
system: summarisationPrompt.system({ tone: 'technical', audience: 'software engineers' }),
messages: [{
role: 'user',
content: summarisationPrompt.user({ content, maxWords: 200, tone: 'technical', audience: 'software engineers' }),
}],
});
Typed prompt builders also make prompts testable. You can unit-test that the builder produces the expected string for a given set of inputs - catching prompt regressions before they reach the LLM call.
When an AI feature is critical to the user experience, a single-provider dependency is a reliability risk. A fallback chain tries the primary provider and falls back to a secondary on failure:
class FallbackProvider implements LLMProvider {
constructor(private providers: LLMProvider[]) {}
async complete(request: LLMRequest): Promise<LLMResponse> {
const errors: Error[] = [];
for (const provider of this.providers) {
try {
return await provider.complete(request);
} catch (err) {
errors.push(err as Error);
// Only fall through for transient errors (5xx, timeout)
// Do not fall through for auth errors (4xx) or input errors
if ((err as any).status && (err as any).status < 500) throw err;
}
}
throw new AggregateError(errors'All providers failed');
}
async *stream(request: LLMRequest): AsyncIterable<string> {
yield* this.providers[0].stream(request); // Stream from primary only
}
}
AI application configuration - model selection, token limits, temperature, retry counts - should never be hardcoded. Use a typed configuration object loaded from environment variables or a config service:
const aiConfig = {
models: {
fast: process.env.AI_FAST_MODEL ?? 'claude-haiku-4-5-20251001',
balanced: process.env.AI_BALANCED_MODEL ?? 'claude-sonnet-5',
capable: process.env.AI_CAPABLE_MODEL ?? 'claude-opus-5',
},
limits: {
maxInputTokens: parseInt(process.env.AI_MAX_INPUT_TOKENS ?? '4000'),
maxOutputTokens: parseInt(process.env.AI_MAX_OUTPUT_TOKENS ?? '2048'),
dailyCostCeilingUsd: parseFloat(process.env.AI_DAILY_COST_CEILING ?? '50'),
},
retry: {
maxAttempts: parseInt(process.env.AI_RETRY_ATTEMPTS ?? '3'),
baseDelayMs: parseInt(process.env.AI_RETRY_DELAY_MS ?? '500'),
},
} as const;
Environment-driven configuration means changing the model for a feature is a deployment configuration change, not a code change - and code changes for configuration are the wrong kind of friction when you are iterating on AI behaviour. The same configuration can be overridden per-environment: Haiku in development, Sonnet in staging, the appropriate model per feature in production.
The abstraction patterns here - provider interfaces, middleware chains, typed prompts, environment configuration - are the same patterns that make conventional TypeScript applications maintainable. AI applications are not architecturally special; they just have a new kind of external dependency that is less predictable than a database or a REST API. Applying the same discipline to that dependency is what separates AI features that survive their first production incident from those that require a rewrite.