Quick Start
Enable distributed mode with Redis:Architecture
The HA module has three components that work together:Heartbeat Service
Each pod writes a heartbeat key to Redis every 10 seconds with a 30-second TTL. When a pod dies, its key expires automatically.Session Takeover
When a request arrives for a session owned by a dead pod, the receiving pod runs an atomic Lua script against Redis:- Read the session data
- Verify the current
nodeIdmatchesexpectedOldNodeId(the dead pod) - Atomically update
nodeIdto the new pod and setreassignedAt/reassignedFrom - Return success or failure
Notification Relay
MCP notifications (progress updates, resource changes) targeting sessions on other pods are relayed via Redis Pub/Sub:- Each pod subscribes to
mcp:ha:notify:{nodeId} - When a notification targets a session on another pod, it is published to that pod’s channel
- The receiving pod delivers the notification to the local transport
Configuration
Configure HA settings viafrontmcp.config.ts:
Orphan Session Scanner
In addition to on-demand session takeover (when a request arrives for a dead pod’s session), FrontMCP runs a periodic orphan scanner that proactively detects and claims sessions from dead pods:- Runs every heartbeat interval (default: 10s)
- Compares all session
nodeIdvalues against alive heartbeat keys - Claims orphaned sessions via the same atomic Lua CAS script
- Fires a callback for each claimed session (logged at INFO level)
transport.persistence.redis) — no other configuration needed. It scans the same key prefix the session store writes to (<keyPrefix>session:, default mcp:transport:session:), so a custom keyPrefix is honoured. The first scan runs after one full heartbeat interval plus the grace period to allow the cluster to stabilize.
Redis Connection
The heartbeat, takeover and transport bus share one dedicated ioredis client built from your top-levelredis config (host, port, password, db, tls, or url). It reconnects automatically, logs connection errors at a rate-limited interval instead of on every retry, and is closed on shutdown. Vercel KV (provider: 'vercel-kv') is REST-only and cannot provide the Lua/pub-sub primitives HA needs, so distributed mode starts without HA when it is configured.
Session TTL
Persisted sessions expire after the first of these that is set:transport.persistence.defaultTtlMstransport.persistence.redis.defaultTtlMs- 1 hour (
3600000)
Redis Unavailable at Startup
If Redis cannot be reached while the server starts, session persistence is disabled for the moment and the server starts. The connection is then retried in the background with exponential backoff (1s doubling to 30s) until it succeeds, after which sessions are persisted and/readyz reports ready. No restart is needed.
Load Balancer Affinity
FrontMCP sets two identifiers on both Streamable HTTP and SSE responses for load balancer routing:- Cookie:
__frontmcp_node--- set during the initialize handshake - Header:
X-FrontMCP-Machine-Id--- set on every response in distributed mode: initialize, message POSTs, DELETE, stateless requests, and SSE. It is applied by a hookableapplyNodeHeadersstage at the start of the streamable-HTTP and stateless flows, so plugins can observe or wrap it.
NGINX Sticky Sessions
The affinity cookie ensures subsequent requests from the same MCP client hit the same pod. If the pod dies, the load balancer routes to a different pod, which triggers session takeover.
SSE-Specific Routing
SSE (Server-Sent Events) requires special attention because the/sse endpoint creates a long-lived connection. POST requests to /message must reach the pod with the active SSE stream.
FrontMCP handles this in two layers:
- LB Affinity (primary): The
__frontmcp_nodecookie is set during SSE initialization, so the load balancer routes subsequent POST requests to the correct pod. - Notification Relay (fallback): If a POST arrives at the wrong pod, FrontMCP detects that the session exists on another node and relays the message via Redis Pub/Sub to the owning pod, which delivers it through the active SSE stream.
Kubernetes
Deploy with 3 replicas and a Redis instance:Errors
Verifying HA
Related
Transport Security
CORS, bind address, DNS rebinding, and host validation
Health Checks
Configure /healthz and /readyz probes
Redis Setup
Redis connection and session store configuration
Runtime Modes
Standalone, distributed, and serverless modes