AI Agents
Test MCP servers with the TypeScript SDK v2: MCP Inspector checks, in-process tests with createMcpHandler, stdio smoke tests and model-facing checks.
Your MCP server's handlers have full unit test coverage, and it still fails the first time Claude connects. A debug console.log corrupted the stdio stream, the schema says orderId while the handler reads order_id, or the tool description never says when to use it. These are protocol and model-facing bugs that handler tests cannot see. MCP server testing means testing the server the way a client uses it.
Test an MCP server at four layers: pure business logic with ordinary unit tests, the protocol surface with an in-process client that lists and calls tools, the real transport with a smoke test that boots the built server, and model behaviour with a small set of prompts that checks which tool Claude picks. Each layer catches bugs the others miss.
few, slow, costs tokens
[ model-facing checks ] does Claude pick the right tool?
[ stdio smoke test ] does the built server boot and answer?
[ in-process protocol ] schemas, results, isError paths
[ unit tests ] business logic, no MCP at all
many, fast, free
The examples below use the MCP TypeScript SDK v2, which splits into @modelcontextprotocol/server and @modelcontextprotocol/client and uses Zod v4 schemas. If you are still on the single @modelcontextprotocol/sdk v1 package, the same layers apply with that package's import paths.
Two structural choices make everything else easy. Build the server in a factory function instead of a module-level singleton, so every test gets a fresh instance. And pass external dependencies into that factory, so tests can swap the database for a fake without module mocking:
// src/server.ts
export type Deps = { findOrder: (id: string) => Promise<Order | null> };
export function createServer(deps: Deps = { findOrder: dbFindOrder }): McpServer {
const server = new McpServer({ name: 'orders', version: '1.0.0' });
server.registerTool(
'get-order',
{
description: 'Look up one order by its ID. Use it when the user mentions an order number such as A1001.',
inputSchema: z.object({ orderId: z.string().min(2).describe('Order ID, for example A1001') }),
outputSchema: z.object({ status: z.string(), total: z.number() }),
},
async ({ orderId }) => {
const order = await deps.findOrder(orderId);
if (!order) {
return {
content: [{ type: 'text', text: 'No order ' + orderId + '. Ask the user to check the order number.' }],
isError: true,
};
}
return {
content: [{ type: 'text', text: 'Order ' + orderId + ': ' + order.status + ', total ' + order.total }],
structuredContent: { status: order.status, total: order.total },
};
},
);
return server;
}
// src/index.ts
void serveStdio(() => createServer());
console.error('orders MCP server running on stdio'); // stderr: stdout carries the protocol
Before writing automated tests, look at what your server actually advertises. The MCP Inspector starts the server, connects as a client and shows every tool, resource and prompt with its schema. Its CLI mode runs the same operations from a terminal or a CI job:
# Interactive: opens the Inspector UI in your browser
npx @modelcontextprotocol/inspector npx tsx src/index.ts
# Scriptable: list tools, then call one
npx @modelcontextprotocol/inspector --cli npx tsx src/index.ts --method tools/list
npx @modelcontextprotocol/inspector --cli npx tsx src/index.ts --method tools/call --tool-name get-order --tool-arg orderId=A1001
Read the generated JSON Schema carefully the first time. It is exactly what the model will see, and it is where mismatches such as a field marked optional that your handler treats as required become obvious.
The SDK v2 can serve a server factory in-process through createMcpHandler, the approach its testing guide recommends. Point a StreamableHTTPClientTransport at the handler's fetch function and the client never opens a socket: each test runs the full protocol, including schema validation and result serialization, in a few milliseconds. With vitest:
// test/server.test.ts
const ORDERS: Record<string, { status: string; total: number }> = {
A1001: { status: 'shipped', total: 42 },
};
describe('orders MCP server', () => {
let handler: ReturnType<typeof createMcpHandler>;
let client: Client;
beforeEach(async () => {
// Fake data source injected through the factory: no database in these tests.
handler = createMcpHandler(() => createServer({ findOrder: async (id) => ORDERS[id] ?? null }));
const transport = new StreamableHTTPClientTransport(new URL('http://test.local/mcp'), {
fetch: (url, init) => handler.fetch(new Request(url, init)), // served in-process, never dialled
});
client = new Client({ name: 'test-harness', version: '1.0.0' }, { versionNegotiation: { mode: 'auto' } });
await client.connect(transport);
});
afterEach(async () => {
await client.close();
await handler.close();
});
it('advertises get-order with a required orderId', async () => {
const { tools } = await client.listTools();
const tool = tools.find((t) => t.name === 'get-order');
expect(tool?.description).toMatch(/order/i);
expect(tool?.inputSchema.required).toContain('orderId');
});
it('returns structured content for a known order', async () => {
const result = await client.callTool({ name: 'get-order', arguments: { orderId: 'A1001' } });
expect(result.isError).toBeFalsy();
expect(result.structuredContent).toEqual({ status: 'shipped', total: 42 });
});
it('reports a missing order as a tool error, not a crash', async () => {
const result = await client.callTool({ name: 'get-order', arguments: { orderId: 'A9999' } });
expect(result.isError).toBe(true);
});
it('refuses input that violates the schema', async () => {
// Pins down how your SDK version surfaces validation failures, so an upgrade cannot change it silently.
const outcome = await client.callTool({ name: 'get-order', arguments: { orderId: 7 } }).catch((e) => e);
expect(outcome instanceof Error || outcome.isError === true).toBe(true);
});
});
A handler failure resolves as an ordinary result with isError: true, not a thrown error, which is exactly what you want: the calling agent sees the message and can recover. Assert on structuredContent for success cases rather than on the text, since the text is for humans and models and will be reworded over time.
Download a ready-made Claude Code subagent that designs, reviews and tests MCP servers, schemas and transports.
Get the MCP server engineer subagentIn-process tests never execute your entry point, your build output or your transport. One smoke test that spawns the compiled server over stdio covers all three, and it is the test that catches the stray console.log:
// test/stdio.smoke.test.ts
it('boots from the built output and lists tools over stdio', async () => {
const transport = new StdioClientTransport({ command: 'node', args: ['dist/index.js'] });
const client = new Client({ name: 'stdio-smoke', version: '1.0.0' });
await client.connect(transport);
try {
const { tools } = await client.listTools();
expect(tools.map((t) => t.name)).toContain('get-order');
} finally {
await client.close();
}
}, 20_000);
Run it after the build step in CI, next to an Inspector CLI call such as --method tools/list. Together they take a few seconds and prove the artifact you ship actually starts.
A server can pass every protocol test and still fail in production because the model misunderstands it. The tool description is a prompt, so test it like one: convert the server's tool list into Claude API tool definitions and check which tool the model chooses for a handful of realistic requests.
import Anthropic from '@anthropic-ai/sdk';
const anthropic = new Anthropic();
// Reuse a connected MCP client from the previous steps.
const { tools } = await client.listTools();
const claudeTools: Anthropic.Tool[] = tools.map((t) => ({
name: t.name,
description: t.description ?? '',
input_schema: t.inputSchema as Anthropic.Tool.InputSchema,
}));
const cases = [
{ prompt: 'Where is my order A1001?', expected: 'get-order' },
{ prompt: 'Has order A2044 shipped yet?', expected: 'get-order' },
{ prompt: 'What are your opening hours?', expected: null },
];
for (const c of cases) {
const res = await anthropic.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
tools: claudeTools,
messages: [{ role: 'user', content: c.prompt }],
});
const call = res.content.find((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use');
const got = call?.name ?? null;
console.log(got === c.expected ? 'PASS' : 'FAIL', JSON.stringify(c.prompt), 'called', got);
}
Keep this suite small, 10 to 30 prompts including a few that should not trigger any tool, and run it when descriptions or schemas change rather than on every commit, because it costs tokens and its results vary slightly between runs. The negative cases matter as much as the positive ones: a tool that fires on unrelated questions is as broken as one that never fires.
orderId, an older handler reads order_id. Protocol tests with real arguments catch it immediately.isError results with an actionable message, as described in tool result error handling.If your server sits behind OAuth or API keys, add tests for rejected and expired credentials too; MCP server authentication covers those flows. For the server itself, start from building a custom MCP server, and for tools versus resources and prompts, MCP resources and prompts.
Use the MCP Inspector for manual checks, and the SDK client in your test runner for automated ones. With the TypeScript SDK v2 you can serve the server in-process through createMcpHandler and point a StreamableHTTPClientTransport at its fetch function, so tests exercise the real protocol without opening a port.
The MCP Inspector is the official debugging tool for Model Context Protocol servers. It starts your server, connects as a client, and lets you list and call tools, resources and prompts from a browser UI. Its CLI mode runs the same operations from the command line, which makes it usable in CI.
Return a result with isError set to true for failures the model should see and react to, such as a missing record or invalid business input. Protocol-level errors are for problems with the request itself. Write a test for each case, because the distinction decides whether the calling agent can recover.
The most common cause is writing logs to stdout: over stdio, stdout carries the JSON-RPC stream, so any console.log corrupts it. Other causes are a build step that was not run, a missing runtime dependency, or a relative path that only resolves from your project directory. A stdio smoke test against the built output catches all of these.