AI & Development

How to Build an AI Content Pipeline with Claude (2026)

How to build a production AI content pipeline - brief intake, Claude generation, quality scoring, and output routing.

Calling Claude in a loop is not a content pipeline. It is a loop. A real content pipeline has guardrails: structured inputs that prevent malformed briefs from reaching the model, a generation stage with quality constraints baked into the prompt, a scoring step that automatically catches outputs below your standard, and an output layer that routes results to the right destination without manual review for every item. The difference matters at scale - running 500 API calls with no quality gate produces 500 outputs of unknown quality. Running 500 calls through a four-stage pipeline produces a set of verified outputs and a separate review queue for the ones that need human attention. This tutorial builds the latter.

AI content pipeline
An AI content pipeline is an automated system that ingests structured content briefs, passes them through an LLM generation stage with constrained prompts, evaluates each output against quality criteria using a second model call, and routes results to storage or a publishing system - enabling production-scale content generation with measurable quality control rather than per-item human review.

What we're building

Content brief (topic, type, constraints)
          │
          ▼
┌─────────────────────────────────────┐
│  Stage 1 - Intake & Validation      │
│  Validates brief schema             │
│  Rejects malformed inputs early     │
└──────────────┬──────────────────────┘
               │
               ▼
┌─────────────────────────────────────┐
│  Stage 2 - Generation               │
│  claude-sonnet-5 (long-form)        │
│  claude-haiku-4-5-20251001 (short)  │
│  Forced tool use → typed output     │
└──────────────┬──────────────────────┘
               │
               ▼
┌─────────────────────────────────────┐
│  Stage 3 - Quality Scoring          │
│  claude-haiku-4-5-20251001 judge    │
│  Scores: accuracy, clarity, length  │
│  Routes: auto-approve vs review     │
└──────────────┬──────────────────────┘
               │
          ┌────┴────┐
          ▼         ▼
      approved    review
      → publish   → queue

Step 1 - Define the brief schema

Every item entering the pipeline must conform to a typed schema. Validating at intake - before any API call - prevents garbage-in-garbage-out and makes the generation prompt predictable:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY });

// Brief schema - validated before any API call
const ContentBriefSchema = z.object({
  id: z.string(),
  type: z.enum(['blog_post', 'product_description', 'faq_answer', 'social_post']),
  topic: z.string().min(10).max(500),
  target_audience: z.string(),
  tone: z.enum(['professional', 'conversational', 'technical', 'persuasive']),
  word_count: z.object({
    min: z.number().int().positive(),
    max: z.number().int().positive(),
  }),
  keywords: z.array(z.string()).max(5),
  constraints: z.array(z.string()).optional(),
});

type ContentBrief = z.infer;

function validateBrief(raw: unknown): { valid: true; brief: ContentBrief } | { valid: false; error: string } {
  const result = ContentBriefSchema.safeParse(raw);
  if (!result.success) {
    return { valid: false, error: result.error.issues.map(i => i.message).join('; ') };
  }
  return { valid: true, brief: result.data };
}

// Typed output schema - what each generated piece must contain
interface GeneratedContent {
  title: string;
  body: string;
  word_count: number;
  keywords_used: string[];
  meta_description: string;
}

Step 2 - Generation stage

Use forced tool use to guarantee the output matches your typed schema. Free-form text generation requires parsing; tool use gives you a structured object directly:

function buildGenerationTool(): Anthropic.Tool {
  return {
    name: 'submit_content',
    description: 'Submit the generated content piece. Call this when the content is complete and ready.',
    input_schema: {
      type: 'object' as const,
      properties: {
        title: { type: 'string', description: 'Content title. Specific and keyword-rich.' },
        body: { type: 'string', description: 'Full content body in plain text.' },
        word_count: { type: 'number', description: 'Actual word count of the body.' },
        keywords_used: {
          type: 'array',
          items: { type: 'string' },
          description: 'Which of the requested keywords appear in the body.',
        },
        meta_description: { type: 'string', description: '130-158 character meta description.' },
      },
      required: ['title', 'body', 'word_count', 'keywords_used', 'meta_description'],
    },
  };
}

function buildSystemPrompt(brief: ContentBrief): string {
  const constraintList = brief.constraints?.map(c => '- ' + c).join('
') ?? 'None';
  return [
    'You are a professional content writer. Produce content that exactly meets the brief.',
    '',
    'Brief:',
    '- Type: ' + brief.type,
    '- Topic: ' + brief.topic,
    '- Audience: ' + brief.target_audience,
    '- Tone: ' + brief.tone,
    '- Word count: ' + brief.word_count.min + '-' + brief.word_count.max + ' words',
    '- Keywords to include: ' + brief.keywords.join(''),
    '',
    'Constraints:',
    constraintList,
    '',
    'Rules:',
    '- Hit the word count range exactly',
    '- Include all provided keywords naturally (not stuffed)',
    '- Do not use filler phrases: "In conclusion", "It is worth noting", "leverage", "seamlessly"',
    '- Submit the finished content using the submit_content tool',
  ].join('
');
}

async function generateContent(brief: ContentBrief): Promise {
  // Select model based on content type and length
  const model = brief.word_count.max > 500
    ? 'claude-sonnet-5'           // Long-form: better quality
    : 'claude-haiku-4-5-20251001'; // Short-form: 5x cheaper

  const response = await client.messages.create({
    model,
    max_tokens: Math.min(8192, brief.word_count.max * 6), // ~6 tokens per word
    system: buildSystemPrompt(brief),
    tools: [buildGenerationTool()],
    tool_choice: { type: 'tool', name: 'submit_content' },
    messages: [{ role: 'user', content: 'Generate the content piece for this brief.' }],
  });

  const toolCall = response.content.find(
    (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
  );
  if (!toolCall) throw new Error('Generation stage: submit_content tool was not called');

  return toolCall.input as GeneratedContent;
}

Step 3 - Quality scoring

A second, cheaper model call scores the output before it leaves the pipeline. Use claude-haiku-4-5-20251001 for scoring - it runs in under 2 seconds and costs ~$0.0003 per evaluation, making it practical to score every item:

interface QualityScore {
  overall: number;          // 0-10
  accuracy: number;         // 0-10 - factual correctness
  clarity: number;          // 0-10 - readability
  brief_adherence: number;  // 0-10 - matches the brief
  word_count_ok: boolean;
  keywords_coverage: number; // % of requested keywords used
  issues: string[];         // specific problems found
  recommendation: 'approve' | 'review' | 'reject';
}

const QUALITY_THRESHOLD = {
  auto_approve: 7.5,  // overall >= 7.5 → auto-publish
  review: 5.0,        // 5.0 <= overall < 7.5 → human review
  reject: 0,          // overall < 5.0 → discard, retry
};

async function scoreContent(brief: ContentBrief, content: GeneratedContent): Promise {
  const prompt = [
    'Score this content against the brief. Be strict.',
    '',
    'Brief topic: ' + brief.topic,
    'Required keywords: ' + brief.keywords.join(''),
    'Required word count: ' + brief.word_count.min + '-' + brief.word_count.max,
    'Required tone: ' + brief.tone,
    '',
    'Content title: ' + content.title,
    'Content word count: ' + content.word_count,
    'Keywords used: ' + content.keywords_used.join(''),
    '',
    'Content body (first 500 chars):',
    content.body.slice(0, 500),
    '',
    'Return JSON only matching this schema:',
    '{ "overall": 0-10, "accuracy": 0-10, "clarity": 0-10, "brief_adherence": 0-10,',
    '  "word_count_ok": boolean, "keywords_coverage": 0-100,',
    '  "issues": string[], "recommendation": "approve"|"review"|"reject" }',
  ].join('
');

  const response = await client.messages.create({
    model: 'claude-haiku-4-5-20251001',
    max_tokens: 512,
    messages: [{ role: 'user', content: prompt }],
  });

  const text = response.content[0].type === 'text' ? response.content[0].text : '{}';
  // Strip any markdown code fences before parsing
  const clean = text.replace(/^```[a-z]*
?/m'').replace(/
?```$/m'').trim();
  return JSON.parse(clean) as QualityScore;
}

Step 4 - Output routing

interface PipelineResult {
  brief_id: string;
  content?: GeneratedContent;
  score?: QualityScore;
  status: 'approved' | 'review' | 'rejected' | 'error';
  error?: string;
  cost_usd?: number;
}

async function runPipeline(rawBrief: unknown): Promise {
  // Stage 1 - Intake validation
  const validation = validateBrief(rawBrief);
  if (!validation.valid) {
    return { brief_id: (rawBrief as any)?.id ?? 'unknown', status: 'error', error: validation.error };
  }
  const brief = validation.brief;

  try {
    // Stage 2 - Generation
    const content = await generateContent(brief);

    // Stage 3 - Quality scoring
    const score = await scoreContent(brief, content);

    // Stage 4 - Routing
    let status: PipelineResult['status'];
    if (score.overall >= QUALITY_THRESHOLD.auto_approve && score.recommendation === 'approve') {
      status = 'approved';
      await publishContent(brief.id, content);       // Write to CMS / database
    } else if (score.overall >= QUALITY_THRESHOLD.review) {
      status = 'review';
      await addToReviewQueue(brief.id, content, score); // Flag for human review
    } else {
      status = 'rejected';
      await logRejection(brief.id, content, score);    // Log for analysis
    }

    return { brief_id: brief.id, content, score, status };
  } catch (err) {
    return { brief_id: brief.id, status: 'error', error: (err as Error).message };
  }
}

Step 5 - Batch processing at scale

async function runBatch(
  briefs: unknown[],
  options: { concurrency: number; onProgress?: (done: number, total: number) => void }
): Promise {
  const results: PipelineResult[] = [];
  const queue = [...briefs];
  let done = 0;

  async function worker(): Promise {
    while (queue.length > 0) {
      const brief = queue.shift();
      if (!brief) break;
      const result = await runPipeline(brief);
      results.push(result);
      done++;
      options.onProgress?.(done, briefs.length);
    }
  }

  await Promise.all(Array.from({ length: options.concurrency }, () => worker()));

  // Summary
  const approved = results.filter(r => r.status === 'approved').length;
  const review   = results.filter(r => r.status === 'review').length;
  const rejected = results.filter(r => r.status === 'rejected').length;
  const errors   = results.filter(r => r.status === 'error').length;

  console.log('Pipeline complete:');
  console.log('  Auto-approved: ' + approved + '/' + briefs.length);
  console.log('  Needs review:  ' + review);
  console.log('  Rejected:      ' + rejected);
  console.log('  Errors:        ' + errors);

  return results;
}

// Run 50 briefs with 5 concurrent workers
const results = await runBatch(briefs, {
  concurrency: 5,
  onProgress: (done, total) => console.log(done + '/' + total),
});

Real cost numbers (2026 pricing)

  • Short-form content (150-300 words, Haiku generation + Haiku scoring): ~$0.001 per piece at scale
  • Long-form content (800-1,500 words, Sonnet generation + Haiku scoring): ~$0.015-0.025 per piece
  • 500 blog posts (long-form): ~$10-15 total API cost
  • 10,000 product descriptions (short-form): ~$10 total

Failure modes to handle

  • Generation timeout - long-form generation can take 30-60 seconds for Sonnet. Set AbortSignal.timeout(90_000) on the API call and retry once before failing the item.
  • Quality score drift - the scoring model may calibrate differently across batches. Spot-check 5% of auto-approved items manually to catch systematic scoring bias.
  • Keyword stuffing - if the model hits keyword targets by repeating them unnaturally, add "Keywords must appear naturally - do not repeat any keyword more than twice" to the system prompt.
  • Rate limits - with 5 concurrent workers and Sonnet, you will hit Tier 1 rate limits quickly. Monitor the x-ratelimit-remaining-tokens response header and back off when it drops below 10,000.

For tracking cost per batch run and identifying which content types consume the most budget, see agent cost management strategies. For building a data extraction variant of this pipeline that pulls structured fields from existing documents rather than generating new content, see the data extraction agent tutorial.