AI Agents

MCP Server Testing: Inspector, Unit and Integration Tests

Test MCP servers with the TypeScript SDK v2: MCP Inspector checks, in-process tests with createMcpHandler, stdio smoke tests and model-facing checks.

Your MCP server's handlers have full unit test coverage, and it still fails the first time Claude connects. A debug console.log corrupted the stdio stream, the schema says orderId while the handler reads order_id, or the tool description never says when to use it. These are protocol and model-facing bugs that handler tests cannot see. MCP server testing means testing the server the way a client uses it.

MCP server testing
MCP server testing is the practice of verifying a Model Context Protocol server at several layers: its business logic, the tools and schemas it advertises over the protocol, the transport it runs on such as stdio, and whether a language model given its tool descriptions calls the right tools with valid arguments.

What to test in an MCP server

Test an MCP server at four layers: pure business logic with ordinary unit tests, the protocol surface with an in-process client that lists and calls tools, the real transport with a smoke test that boots the built server, and model behaviour with a small set of prompts that checks which tool Claude picks. Each layer catches bugs the others miss.

                   few, slow, costs tokens
  [ model-facing checks   ]  does Claude pick the right tool?
  [ stdio smoke test      ]  does the built server boot and answer?
  [ in-process protocol   ]  schemas, results, isError paths
  [ unit tests            ]  business logic, no MCP at all
                   many, fast, free

The examples below use the MCP TypeScript SDK v2, which splits into @modelcontextprotocol/server and @modelcontextprotocol/client and uses Zod v4 schemas. If you are still on the single @modelcontextprotocol/sdk v1 package, the same layers apply with that package's import paths.

Step 1: structure the server for testing

Two structural choices make everything else easy. Build the server in a factory function instead of a module-level singleton, so every test gets a fresh instance. And pass external dependencies into that factory, so tests can swap the database for a fake without module mocking:

// src/server.ts



export type Deps = { findOrder: (id: string) => Promise<Order | null> };

export function createServer(deps: Deps = { findOrder: dbFindOrder }): McpServer {
  const server = new McpServer({ name: 'orders', version: '1.0.0' });

  server.registerTool(
    'get-order',
    {
      description: 'Look up one order by its ID. Use it when the user mentions an order number such as A1001.',
      inputSchema: z.object({ orderId: z.string().min(2).describe('Order ID, for example A1001') }),
      outputSchema: z.object({ status: z.string(), total: z.number() }),
    },
    async ({ orderId }) => {
      const order = await deps.findOrder(orderId);
      if (!order) {
        return {
          content: [{ type: 'text', text: 'No order ' + orderId + '. Ask the user to check the order number.' }],
          isError: true,
        };
      }
      return {
        content: [{ type: 'text', text: 'Order ' + orderId + ': ' + order.status + ', total ' + order.total }],
        structuredContent: { status: order.status, total: order.total },
      };
    },
  );
  return server;
}

// src/index.ts


void serveStdio(() => createServer());
console.error('orders MCP server running on stdio'); // stderr: stdout carries the protocol

Step 2: poke it by hand with the MCP Inspector

Before writing automated tests, look at what your server actually advertises. The MCP Inspector starts the server, connects as a client and shows every tool, resource and prompt with its schema. Its CLI mode runs the same operations from a terminal or a CI job:

# Interactive: opens the Inspector UI in your browser
npx @modelcontextprotocol/inspector npx tsx src/index.ts

# Scriptable: list tools, then call one
npx @modelcontextprotocol/inspector --cli npx tsx src/index.ts --method tools/list
npx @modelcontextprotocol/inspector --cli npx tsx src/index.ts --method tools/call --tool-name get-order --tool-arg orderId=A1001

Read the generated JSON Schema carefully the first time. It is exactly what the model will see, and it is where mismatches such as a field marked optional that your handler treats as required become obvious.

Step 3: in-process protocol tests

The SDK v2 can serve a server factory in-process through createMcpHandler, the approach its testing guide recommends. Point a StreamableHTTPClientTransport at the handler's fetch function and the client never opens a socket: each test runs the full protocol, including schema validation and result serialization, in a few milliseconds. With vitest:

// test/server.test.ts




const ORDERS: Record<string, { status: string; total: number }> = {
  A1001: { status: 'shipped', total: 42 },
};

describe('orders MCP server', () => {
  let handler: ReturnType<typeof createMcpHandler>;
  let client: Client;

  beforeEach(async () => {
    // Fake data source injected through the factory: no database in these tests.
    handler = createMcpHandler(() => createServer({ findOrder: async (id) => ORDERS[id] ?? null }));
    const transport = new StreamableHTTPClientTransport(new URL('http://test.local/mcp'), {
      fetch: (url, init) => handler.fetch(new Request(url, init)), // served in-process, never dialled
    });
    client = new Client({ name: 'test-harness', version: '1.0.0' }, { versionNegotiation: { mode: 'auto' } });
    await client.connect(transport);
  });

  afterEach(async () => {
    await client.close();
    await handler.close();
  });

  it('advertises get-order with a required orderId', async () => {
    const { tools } = await client.listTools();
    const tool = tools.find((t) => t.name === 'get-order');
    expect(tool?.description).toMatch(/order/i);
    expect(tool?.inputSchema.required).toContain('orderId');
  });

  it('returns structured content for a known order', async () => {
    const result = await client.callTool({ name: 'get-order', arguments: { orderId: 'A1001' } });
    expect(result.isError).toBeFalsy();
    expect(result.structuredContent).toEqual({ status: 'shipped', total: 42 });
  });

  it('reports a missing order as a tool error, not a crash', async () => {
    const result = await client.callTool({ name: 'get-order', arguments: { orderId: 'A9999' } });
    expect(result.isError).toBe(true);
  });

  it('refuses input that violates the schema', async () => {
    // Pins down how your SDK version surfaces validation failures, so an upgrade cannot change it silently.
    const outcome = await client.callTool({ name: 'get-order', arguments: { orderId: 7 } }).catch((e) => e);
    expect(outcome instanceof Error || outcome.isError === true).toBe(true);
  });
});

A handler failure resolves as an ordinary result with isError: true, not a thrown error, which is exactly what you want: the calling agent sees the message and can recover. Assert on structuredContent for success cases rather than on the text, since the text is for humans and models and will be reworded over time.

Building MCP servers regularly?

Download a ready-made Claude Code subagent that designs, reviews and tests MCP servers, schemas and transports.

Get the MCP server engineer subagent

Step 4: a stdio smoke test against the build

In-process tests never execute your entry point, your build output or your transport. One smoke test that spawns the compiled server over stdio covers all three, and it is the test that catches the stray console.log:

// test/stdio.smoke.test.ts



it('boots from the built output and lists tools over stdio', async () => {
  const transport = new StdioClientTransport({ command: 'node', args: ['dist/index.js'] });
  const client = new Client({ name: 'stdio-smoke', version: '1.0.0' });
  await client.connect(transport);
  try {
    const { tools } = await client.listTools();
    expect(tools.map((t) => t.name)).toContain('get-order');
  } finally {
    await client.close();
  }
}, 20_000);

Run it after the build step in CI, next to an Inspector CLI call such as --method tools/list. Together they take a few seconds and prove the artifact you ship actually starts.

Step 5: check how a model uses your tools

A server can pass every protocol test and still fail in production because the model misunderstands it. The tool description is a prompt, so test it like one: convert the server's tool list into Claude API tool definitions and check which tool the model chooses for a handful of realistic requests.

import Anthropic from '@anthropic-ai/sdk';

const anthropic = new Anthropic();

// Reuse a connected MCP client from the previous steps.
const { tools } = await client.listTools();
const claudeTools: Anthropic.Tool[] = tools.map((t) => ({
  name: t.name,
  description: t.description ?? '',
  input_schema: t.inputSchema as Anthropic.Tool.InputSchema,
}));

const cases = [
  { prompt: 'Where is my order A1001?', expected: 'get-order' },
  { prompt: 'Has order A2044 shipped yet?', expected: 'get-order' },
  { prompt: 'What are your opening hours?', expected: null },
];

for (const c of cases) {
  const res = await anthropic.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 1024,
    tools: claudeTools,
    messages: [{ role: 'user', content: c.prompt }],
  });
  const call = res.content.find((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use');
  const got = call?.name ?? null;
  console.log(got === c.expected ? 'PASS' : 'FAIL', JSON.stringify(c.prompt), 'called', got);
}

Keep this suite small, 10 to 30 prompts including a few that should not trigger any tool, and run it when descriptions or schemas change rather than on every commit, because it costs tokens and its results vary slightly between runs. The negative cases matter as much as the positive ones: a tool that fires on unrelated questions is as broken as one that never fires.

Failure modes these tests catch

  • stdout pollution. Any write to stdout in a stdio server corrupts the JSON-RPC stream. Only the smoke test sees it.
  • Schema and handler drift. The schema says orderId, an older handler reads order_id. Protocol tests with real arguments catch it immediately.
  • Crashes instead of tool errors. A thrown exception for "not found" gives the model nothing useful to act on. Assert that expected failures come back as isError results with an actionable message, as described in tool result error handling.
  • Shared state between tests. A module-level server or cache makes test order matter. The factory pattern removes it.
  • Vague descriptions. Only the model-facing check shows that "Get order" is too thin and "Use it when the user mentions an order number" is not.

If your server sits behind OAuth or API keys, add tests for rejected and expired credentials too; MCP server authentication covers those flows. For the server itself, start from building a custom MCP server, and for tools versus resources and prompts, MCP resources and prompts.

FAQ

How do I test an MCP server without a client app?

Use the MCP Inspector for manual checks, and the SDK client in your test runner for automated ones. With the TypeScript SDK v2 you can serve the server in-process through createMcpHandler and point a StreamableHTTPClientTransport at its fetch function, so tests exercise the real protocol without opening a port.

What is the MCP Inspector?

The MCP Inspector is the official debugging tool for Model Context Protocol servers. It starts your server, connects as a client, and lets you list and call tools, resources and prompts from a browser UI. Its CLI mode runs the same operations from the command line, which makes it usable in CI.

Should an MCP tool throw or return isError?

Return a result with isError set to true for failures the model should see and react to, such as a missing record or invalid business input. Protocol-level errors are for problems with the request itself. Write a test for each case, because the distinction decides whether the calling agent can recover.

Why does my MCP server work in tests but fail in Claude Desktop or Claude Code?

The most common cause is writing logs to stdout: over stdio, stdout carries the JSON-RPC stream, so any console.log corrupts it. Other causes are a build step that was not run, a missing runtime dependency, or a relative path that only resolves from your project directory. A stdio smoke test against the built output catches all of these.