AI Agents
AI agents fail in ways that standard logs cannot capture. How to instrument agentic loops with structured traces, span tracking, tool call logging.
When an agent in production returns a wrong answer or fails to complete, your standard application logs will show you that it happened. They will not show you which of the twelve tool calls in that run returned unexpected data, which iteration the loop diverged from the expected path, or how much the failed run cost before it gave up. Debugging agentic systems without traces is like debugging a database without query logs - you can tell something went wrong, but finding the cause requires guesswork. Structured agent tracing converts guesswork into inspection.
Agent observability is the ability to reconstruct exactly what happened in an agentic run after it completes: every model call, every tool invocation, every tool result, the token counts and costs at each step, the full message history at each iteration, and the final output or failure reason. With this data, debugging is a read operation, not an investigation.
Model every agentic run as a trace with a root span and child spans per iteration. The root span captures the overall run; each iteration is a child span with its own model call, tool calls, and results:
interface AgentSpan {
spanId: string;
parentSpanId?: string;
type: 'run' | 'iteration' | 'model_call' | 'tool_call';
startedAt: string; // ISO timestamp
completedAt?: string;
durationMs?: number;
status: 'running' | 'completed' | 'failed' | 'timeout';
error?: string;
}
interface RunSpan extends AgentSpan {
type: 'run';
runId: string;
goal: string;
model: string;
totalIterations: number;
totalInputTokens: number;
totalOutputTokens: number;
estimatedCostUsd: number;
finalOutput?: string;
}
interface IterationSpan extends AgentSpan {
type: 'iteration';
iterationNumber: number;
stopReason: string;
inputTokens: number;
outputTokens: number;
toolCallsRequested: string[];
}
interface ToolCallSpan extends AgentSpan {
type: 'tool_call';
toolName: string;
toolInput: unknown;
toolOutput?: string;
outputTruncated: boolean;
isError: boolean;
}
Wrap the agent loop with a tracer that creates spans automatically. The tracer does not change the loop's logic - it only records what happens alongside it:
import Anthropic from '@anthropic-ai/sdk';
class AgentTracer {
private spans: AgentSpan[] = [];
private runId = randomUUID();
startRun(goal: string, model: string): RunSpan {
const span: RunSpan = {
spanId: this.runId,
type: 'run',
runId: this.runId,
goal,
model,
startedAt: new Date().toISOString(),
status: 'running',
totalIterations: 0,
totalInputTokens: 0,
totalOutputTokens: 0,
estimatedCostUsd: 0,
};
this.spans.push(span);
return span;
}
recordIteration(
iterationNumber: number,
response: Anthropic.Message,
runSpan: RunSpan
): IterationSpan {
const span: IterationSpan = {
spanId: randomUUID(),
parentSpanId: runSpan.spanId,
type: 'iteration',
iterationNumber,
startedAt: new Date().toISOString(),
completedAt: new Date().toISOString(),
status: 'completed',
stopReason: response.stop_reason ?? 'unknown',
inputTokens: response.usage.input_tokens,
outputTokens: response.usage.output_tokens,
toolCallsRequested: response.content
.filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use')
.map(b => b.name),
durationMs: 0,
};
// Accumulate into run span
runSpan.totalIterations++;
runSpan.totalInputTokens += response.usage.input_tokens;
runSpan.totalOutputTokens += response.usage.output_tokens;
runSpan.estimatedCostUsd += calculateCost(
response.model,
response.usage.input_tokens,
response.usage.output_tokens
);
this.spans.push(span);
return span;
}
recordToolCall(
iterationSpan: IterationSpan,
toolName: string,
toolInput: unknown,
toolOutput: string,
isError: boolean,
startedAt: Date
): ToolCallSpan {
const MAX_OUTPUT_LENGTH = 2000;
const span: ToolCallSpan = {
spanId: randomUUID(),
parentSpanId: iterationSpan.spanId,
type: 'tool_call',
toolName,
toolInput,
toolOutput: toolOutput.slice(0, MAX_OUTPUT_LENGTH),
outputTruncated: toolOutput.length > MAX_OUTPUT_LENGTH,
isError,
startedAt: startedAt.toISOString(),
completedAt: new Date().toISOString(),
durationMs: Date.now() - startedAt.getTime(),
status: isError ? 'failed' : 'completed',
};
this.spans.push(span);
return span;
}
finalise(runSpan: RunSpan, output?: string, error?: string): void {
runSpan.status = error ? 'failed' : 'completed';
runSpan.completedAt = new Date().toISOString();
runSpan.finalOutput = output;
runSpan.error = error;
this.flush(runSpan);
}
private flush(runSpan: RunSpan): void {
// Ship to your observability backend
console.log(JSON.stringify({ runId: runSpan.runId, spans: this.spans }));
}
}
function calculateCost(model: string, inputTokens: number, outputTokens: number): number {
const pricing: Record = {
// Per million tokens
'claude-opus-5': { input: 15.00, output: 75.00 },
'claude-sonnet-5': { input: 3.00, output: 15.00 },
'claude-haiku-4-5-20251001': { input: 0.80, output: 4.00 },
};
const modelPricing = pricing[model] ?? pricing['claude-sonnet-5'];
return (inputTokens / 1_000_000 * modelPricing.input) +
(outputTokens / 1_000_000 * modelPricing.output);
}
For production systems, ship traces to a backend that supports structured queries. The simplest viable setup: write each run's JSON trace to a Firestore document, keyed by runId. More scalable: a time-series compatible store like BigQuery or a dedicated LLM observability platform (Langfuse, Helicone, Braintrust).
The queries you will run most often:
status = 'failed' and startedAt > now - 24hestimatedCostUsd > thresholdtoolName = X AND isError = truetotalIterations over successful runsThree failure patterns that are invisible in standard logs but immediately obvious in traces:
stop_reason: 'end_turn' but the final output does not address the goal. Visible by comparing the goal in the root span against the final output - automatable with a simple LLM-as-judge evaluation on each completed trace.Tracing works alongside cost controls - see agent cost management for model tiering and token budgets that reduce the per-run spend you are now measuring. For the error recovery patterns that traces reveal are needed, see agent error recovery strategies.