AI Agents

The Agentic Loop Explained: How AI Agents Actually Work

The observe-think-act loop is the core of every AI agent. This guide explains how the loop works, when it terminates, how tools fit in.

The word "agent" gets applied to a wide range of systems in 2026 - from a single LLM call with a system prompt to a coordinated network of specialised models. What makes something an agent rather than a chain or a chatbot is the loop: the ability to take an action, observe the result, and decide what to do next based on what it observed. Understanding this loop in precise terms is foundational to debugging agent behaviour, designing agent architectures, and reasoning about where agents will succeed and where they will fail.

Agentic loop
An agentic loop is the core execution cycle of an AI agent: the model receives an observation, decides on an action (which may be a tool call or a final answer), the action is executed externally, and the result becomes the next observation - repeating until the model returns stop_reason: 'end_turn' or a maximum iteration limit is reached.

What the agentic loop is - the direct answer

The agentic loop is a repeated cycle of four steps: (1) the agent receives an observation about the current state of its environment, (2) the model processes that observation along with its goal and history to decide on an action, (3) the action is executed - a tool call, a message, a file write - and (4) the result of that action becomes the next observation. The loop continues until the agent decides it has completed its goal, reaches a maximum iteration limit, or is interrupted externally.

In code, the raw loop is simpler than most descriptions make it sound:

import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

async function runAgentLoop(goal: string, tools: Anthropic.Tool[], maxIterations = 20) {
  const messages: Anthropic.MessageParam[] = [
    { role: 'user', content: goal }
  ];

  for (let i = 0; i < maxIterations; i++) {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 4096,
      tools,
      messages,
    });

    // Append the assistant's response to history
    messages.push({ role: 'assistant', content: response.content });

    // If the model stopped naturally (no tool calls), the task is done
    if (response.stop_reason === 'end_turn') {
      return extractFinalAnswer(response.content);
    }

    // If the model wants to use tools, execute them and feed results back
    if (response.stop_reason === 'tool_use') {
      const toolResults = await executeTools(response.content);
      messages.push({ role: 'user', content: toolResults });
      // Loop continues - next iteration processes the tool results
    }
  }

  throw new Error(`Agent did not complete within ${maxIterations} iterations`);
}

The core insight: the message history is the agent's working memory. Every observation, every action, every result accumulates there. The model at each iteration has access to everything that happened before. This is not magic - it is context management.

How tool use fits into the loop

Tools are how the agent takes actions that affect the world outside the model. Without tools, the agent can only produce text - useful, but not agentic. With tools, the agent can read files, call APIs, query databases, write to disk, and trigger external processes.

The mechanics: when the model decides to use a tool, it returns a tool_use content block specifying the tool name and arguments. Your code - not the model - executes the tool and returns the result as a tool_result content block in the next user message. The model never executes code directly. It requests actions; your orchestration layer executes them and reports back.

async function executeTools(
  contentBlocks: Anthropic.ContentBlock[]
): Promise {
  const toolUseBlocks = contentBlocks.filter(
    (b): b is Anthropic.ToolUseBlock => b.type === 'tool_use'
  );

  return Promise.all(
    toolUseBlocks.map(async (block) => {
      try {
        const result = await dispatchTool(block.name, block.input);
        return {
          type: 'tool_result' as const,
          tool_use_id: block.id,
          content: JSON.stringify(result),
        };
      } catch (err) {
        return {
          type: 'tool_result' as const,
          tool_use_id: block.id,
          content: `Error: ${(err as Error).message}`,
          is_error: true,
        };
      }
    })
  );
}

The is_error: true flag in the error case is important - it signals to the model that the tool failed rather than returned an empty result. Models handle explicit errors better than silently empty tool results.

How the loop terminates

Three termination conditions exist, and only one of them is "the agent succeeded":

  • Natural completion (stop_reason: 'end_turn') - The model produced a text response without requesting any tool calls. This is the success path: the agent decided it had enough information to produce a final answer and stopped on its own.
  • Max iteration limit - The loop hit the ceiling you set without a natural completion. This is not a success - it is a signal that the agent is stuck in a cycle, the task is too complex for the given iteration budget, or the goal is underspecified. Log these and investigate the message history.
  • External interruption - Your code or a human operator stopped the loop. This can be intentional (a user cancels the task) or a safety mechanism (a circuit breaker triggered).

The max iteration limit is not a timeout - it is a depth constraint. Set it based on how many tool calls the task genuinely requires. A research task that calls a search API 5 times and then synthesises results needs an iteration budget of at least 7 to 10 (including some back-and-forth). A task with a budget of 3 will fail for anything non-trivial.

The role of the system prompt in the loop

The system prompt sets the context that persists across every iteration of the loop. It is not re-evaluated at each step - it is the fixed context that the model carries throughout. Put in the system prompt: the agent's identity, its goal, the tools available and when to use them, output format requirements, and any hard constraints on behaviour.

Do not put in the system prompt: the current state of the task, user-specific data, or anything that changes between runs. That belongs in the initial user message or in tool results as the loop progresses.

What goes wrong and how to debug it

The most common agentic loop failure modes in production:

  • Infinite tool call loops - The agent calls the same tool repeatedly with the same arguments, getting the same result, never reaching completion. Cause: the tool result is not informative enough for the model to change its plan. Fix: add context to tool results explaining what the result means, not just what it is. "No results found" is better than an empty array; "No results found for query X - consider broadening the search terms" is better still.
  • Context window overflow - Long tool results accumulate in the message history until the total context exceeds the model's limit. Fix: truncate large tool results before returning them (keep the first 2,000 characters of a 50,000-character web page, not the full text). Track total token count across iterations and summarise history when it approaches the limit.
  • Goal drift - The agent starts solving a sub-problem correctly but loses track of the original goal. Cause: the original goal gets buried under many layers of tool results. Fix: repeat the goal in the system prompt and optionally prepend it to tool results ("Working toward: [original goal]. Tool result: ...").
  • Premature termination - The agent stops and returns an answer before it has actually solved the problem. Cause: the model is rewarded by training to avoid tool calls when it believes it can answer from parametric knowledge. Fix: add explicit instructions in the system prompt that the agent must verify claims with tools before finalising.

When the agentic loop is NOT the right choice

The loop has overhead: multiple API calls, accumulated latency, compounding cost per iteration. For tasks with a predictable number of steps and no need for dynamic branching, a fixed prompt chain is simpler, cheaper, and more reliable. Use the agentic loop when the number of steps is unknown in advance, when the agent needs to make decisions based on intermediate results, or when external state (a database, a web page, a file system) must be read and acted upon dynamically. For "process this document and extract these five fields", a single API call with tool use is correct. For "research this topic and compile a structured report", the loop is correct.

Once the agentic loop is stable, the next step is structuring how the loop delegates work to specialised agents - see the supervisor pattern for orchestrator and worker architecture. For production concerns like tracing each loop iteration and tracking per-run cost, see agent tracing and observability.