Maru@maru

Dev Hub

Translated from KoreanView original

Building an MCP Server with Fastify: From API Schema Automation to Redis Scaling

As the Model Context Protocol (MCP) for connecting AI agents to external tools rapidly gains traction, there is a growing need for cloud backend architectures suitable for production—moving beyond local terminal-based stdio environments or single SSE connections. By combining the high-throughput performance of Fastify with the modern MCP development ecosystem, you can design a high-performance agent infrastructure capable of seamless horizontal scaling even under heavy traffic. This guide introduces a robust approach to building MCP servers ready for production, covering everything from automatic tool registration via API schemas to session scaling using Redis.

Streamable HTTP and @modelcontextprotocol/fastify v2.0

Starting with SDK v2.0, which adopts the latest Model Context Protocol specifications, Streamable HTTP has become the standard transport layer, replacing the previous SSE method that was complex and burdensome to maintain. This new standard controls bidirectional stream lifecycles with clients via HTTP POST request-based response streaming, cleanly solving distributed infrastructure operational challenges like firewall traversal and load balancing.

In a Node.js environment, combining the official @modelcontextprotocol/fastify adapter with the @modelcontextprotocol/node package allows you to easily launch production-grade Streamable HTTP endpoints on top of the high-speed Fastify framework. Specifically, to fundamentally block DNS rebinding attacks that can easily occur when running in local host environments, you must implement the hostHeaderValidation middleware.

Below is the standard code pattern for building a secure Streamable HTTP server based on the Fastify v2.0 SDK.

typescript

This pattern is highly advantageous for stateless proxy architectures that handle traffic by connecting a one-time transport for every client request. Since it directly connects Node.js's native IncomingMessage and ServerResponse streams, it can handle various agent JSON-RPC requests stably, lightly, and flexibly.

Automatic MCP Tool Registration via API Schema

Integrating API development and tool specification management for AI agents can significantly increase development productivity. The @mcp-it/fastify plugin leverages Fastify’s unique route compilation lifecycle to automate this process. At server startup, it calls the onRoute hook to discover all API endpoints registered in the application, dynamically extracting unique names and detailed descriptions from each route's schema object to be used as tool identifiers.

The plugin then converts the JSON schema specifications declared as request parameters into an MCP tool format that the agent can understand in real-time. When an agent executes a specific tool via an MCP client, the plugin maps the passed parameter payload into a virtual request object and routes it internally. This allows you to leverage existing controller logic without having to re-implement complex API call mediation logic.

This automated schema extraction approach creates powerful synergy when used in conjunction with the TypeBox validation library in a TypeScript environment. TypeBox schemas declared via @fastify/type-provider-typebox serve simultaneously as TypeScript type checks at compile-time and API validation at runtime. Because this design data is converted directly into MCP tool specifications without modification, updating a single line of code ensures that the API controller, validation layer, and AI agent tool specs remain perfectly synchronized.

Below is a code example of configuring an API endpoint using TypeBox while simultaneously exposing it as an MCP tool.

typescript

With this configuration, when an AI agent calls the route via MCP, the internal Fastify controller logic executes safely through the virtual request interface.

Horizontal Scaling of MCP Sessions with Redis

Typical Model Context Protocol servers store session state and connection information in memory, which creates a structural limitation where scaling instances horizontally in response to traffic causes session disconnections. The @platformatic/mcp plugin developed by the Platformatic team solves this by transparently transitioning session management and message delivery systems to a Redis-based distributed architecture.

This plugin completely alters internal operational processes simply by setting the redis option. Existing in-memory session stores are automatically replaced with RedisSessionStore shared by multiple nodes. Simultaneously, it adopts the distributed Pub/Sub library mqemitter-redis as the backplane, broadcasting all control events delivered to session-unique channels (mcp/session/{sessionId}/message) to all sub-nodes within the cluster in real-time.

typescript

By adopting this architecture, you can maintain seamless agent sessions even if clients are randomly routed to different server instances by a load balancer. Even if a session is disconnected due to temporary network issues, the Last-Event-ID header allows the connection stream to be resumed smoothly from the previous history stored in Redis. In particular, even when asynchronous background tasks are being processed on a specific node, clients can poll for progress across different distributed nodes to check results safely, completing a true enterprise-grade high-availability agent layer.

Architecture Checklist for Production Deployment

When building a high-availability agent backend by combining Fastify with the modern Model Context Protocol ecosystem, you must verify three key design elements. First, enable the host header validation feature provided in the official SDK v2.0 to protect your agent execution environment from security threats such as DNS rebinding. Second, in production environments where concurrent requests are heavy, ensure that you set clear execution time limits so that tool execution speed does not become a bottleneck for the entire backend. Finally, to prevent session loss during horizontal instance scaling, adopt a Redis-based session store and distributed messaging architecture to complete a high-performance agent infrastructure optimized for the cloud.

Reference Links

Loading comments…