AI Agents

Agent Tracing and Observability: Debugging Agentic Runs in Production

AI agents fail in ways that standard logs cannot capture. How to instrument agentic loops with structured traces, span tracking, tool call logging.

When an agent in production returns a wrong answer or fails to complete, your standard application logs will show you that it happened. They will not show you which of the twelve tool calls in that run returned unexpected data, which iteration the loop diverged from the expected path, or how much the failed run cost before it gave up. Debugging agentic systems without traces is like debugging a database without query logs - you can tell something went wrong, but finding the cause requires guesswork. Structured agent tracing converts guesswork into inspection.

Agent trace
An agent trace is a structured log of an agentic run composed of nested spans - a root run span capturing overall metadata (goal, model, total cost), iteration spans per loop cycle, and tool call spans per tool invocation - that together allow post-hoc reconstruction of exactly what the agent did, in what order, and at what cost.

What agent observability is - the core signal

Agent observability is the ability to reconstruct exactly what happened in an agentic run after it completes: every model call, every tool invocation, every tool result, the token counts and costs at each step, the full message history at each iteration, and the final output or failure reason. With this data, debugging is a read operation, not an investigation.

The run trace structure

Model every agentic run as a trace with a root span and child spans per iteration. The root span captures the overall run; each iteration is a child span with its own model call, tool calls, and results:

interface AgentSpan {
  spanId: string;
  parentSpanId?: string;
  type: 'run' | 'iteration' | 'model_call' | 'tool_call';
  startedAt: string;         // ISO timestamp
  completedAt?: string;
  durationMs?: number;
  status: 'running' | 'completed' | 'failed' | 'timeout';
  error?: string;
}

interface RunSpan extends AgentSpan {
  type: 'run';
  runId: string;
  goal: string;
  model: string;
  totalIterations: number;
  totalInputTokens: number;
  totalOutputTokens: number;
  estimatedCostUsd: number;
  finalOutput?: string;
}

interface IterationSpan extends AgentSpan {
  type: 'iteration';
  iterationNumber: number;
  stopReason: string;
  inputTokens: number;
  outputTokens: number;
  toolCallsRequested: string[];
}

interface ToolCallSpan extends AgentSpan {
  type: 'tool_call';
  toolName: string;
  toolInput: unknown;
  toolOutput?: string;
  outputTruncated: boolean;
  isError: boolean;
}

Instrumenting the agent loop

Wrap the agent loop with a tracer that creates spans automatically. The tracer does not change the loop's logic - it only records what happens alongside it:

import Anthropic from '@anthropic-ai/sdk';

class AgentTracer {
  private spans: AgentSpan[] = [];
  private runId = randomUUID();

  startRun(goal: string, model: string): RunSpan {
    const span: RunSpan = {
      spanId: this.runId,
      type: 'run',
      runId: this.runId,
      goal,
      model,
      startedAt: new Date().toISOString(),
      status: 'running',
      totalIterations: 0,
      totalInputTokens: 0,
      totalOutputTokens: 0,
      estimatedCostUsd: 0,
    };
    this.spans.push(span);
    return span;
  }

  recordIteration(
    iterationNumber: number,
    response: Anthropic.Message,
    runSpan: RunSpan
  ): IterationSpan {
    const span: IterationSpan = {
      spanId: randomUUID(),
      parentSpanId: runSpan.spanId,
      type: 'iteration',
      iterationNumber,
      startedAt: new Date().toISOString(),
      completedAt: new Date().toISOString(),
      status: 'completed',
      stopReason: response.stop_reason ?? 'unknown',
      inputTokens: response.usage.input_tokens,
      outputTokens: response.usage.output_tokens,
      toolCallsRequested: response.content
        .filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use')
        .map(b => b.name),
      durationMs: 0,
    };

    // Accumulate into run span
    runSpan.totalIterations++;
    runSpan.totalInputTokens += response.usage.input_tokens;
    runSpan.totalOutputTokens += response.usage.output_tokens;
    runSpan.estimatedCostUsd += calculateCost(
      response.model,
      response.usage.input_tokens,
      response.usage.output_tokens
    );

    this.spans.push(span);
    return span;
  }

  recordToolCall(
    iterationSpan: IterationSpan,
    toolName: string,
    toolInput: unknown,
    toolOutput: string,
    isError: boolean,
    startedAt: Date
  ): ToolCallSpan {
    const MAX_OUTPUT_LENGTH = 2000;
    const span: ToolCallSpan = {
      spanId: randomUUID(),
      parentSpanId: iterationSpan.spanId,
      type: 'tool_call',
      toolName,
      toolInput,
      toolOutput: toolOutput.slice(0, MAX_OUTPUT_LENGTH),
      outputTruncated: toolOutput.length > MAX_OUTPUT_LENGTH,
      isError,
      startedAt: startedAt.toISOString(),
      completedAt: new Date().toISOString(),
      durationMs: Date.now() - startedAt.getTime(),
      status: isError ? 'failed' : 'completed',
    };
    this.spans.push(span);
    return span;
  }

  finalise(runSpan: RunSpan, output?: string, error?: string): void {
    runSpan.status = error ? 'failed' : 'completed';
    runSpan.completedAt = new Date().toISOString();
    runSpan.finalOutput = output;
    runSpan.error = error;
    this.flush(runSpan);
  }

  private flush(runSpan: RunSpan): void {
    // Ship to your observability backend
    console.log(JSON.stringify({ runId: runSpan.runId, spans: this.spans }));
  }
}

Cost calculation per model

function calculateCost(model: string, inputTokens: number, outputTokens: number): number {
  const pricing: Record = {
    // Per million tokens
    'claude-opus-5':                { input: 15.00, output: 75.00 },
    'claude-sonnet-5':              { input: 3.00,  output: 15.00 },
    'claude-haiku-4-5-20251001':    { input: 0.80,  output: 4.00  },
  };

  const modelPricing = pricing[model] ?? pricing['claude-sonnet-5'];
  return (inputTokens / 1_000_000 * modelPricing.input) +
         (outputTokens / 1_000_000 * modelPricing.output);
}

Storing and querying traces

For production systems, ship traces to a backend that supports structured queries. The simplest viable setup: write each run's JSON trace to a Firestore document, keyed by runId. More scalable: a time-series compatible store like BigQuery or a dedicated LLM observability platform (Langfuse, Helicone, Braintrust).

The queries you will run most often:

  • Find all failed runs in the last 24 hours - Filter on status = 'failed' and startedAt > now - 24h
  • Find runs that exceeded cost budget - Filter on estimatedCostUsd > threshold
  • Find runs where a specific tool errored - Query tool_call spans where toolName = X AND isError = true
  • Average iterations per successful run - Aggregate totalIterations over successful runs

The three metrics every agent dashboard needs

  • Success rate - Percentage of runs that completed successfully vs failed or timed out. Segment by agent type. A success rate below 85% in production indicates a systemic issue in tool reliability or goal specification.
  • p95 iteration count - The 95th-percentile number of iterations per run. High p95 with low median indicates a subset of inputs is causing the agent to loop excessively. Look at those traces specifically.
  • Cost per successful run - Total spend divided by successful completions. This tells you the actual unit economics of your agent and whether model or tool optimisations are working.

Failure modes that only traces reveal

Three failure patterns that are invisible in standard logs but immediately obvious in traces:

  • Tool call oscillation - The agent alternates between two tool calls with slightly different arguments, never making progress. Visible as a repeating pattern in the tool_call spans of an iteration sequence.
  • Token leakage - A tool consistently returns 50,000-character results that bloat the context, causing the final iterations to run at massively higher cost. Visible in per-iteration token counts that increase monotonically.
  • Silent wrong answers - The agent completes with stop_reason: 'end_turn' but the final output does not address the goal. Visible by comparing the goal in the root span against the final output - automatable with a simple LLM-as-judge evaluation on each completed trace.

Tracing works alongside cost controls - see agent cost management for model tiering and token budgets that reduce the per-run spend you are now measuring. For the error recovery patterns that traces reveal are needed, see agent error recovery strategies.