Skip to main content
This guide walks through adding production-grade traffic controls to your FrontMCP server using the Guard system.
Prerequisites: You should have a working FrontMCP server with at least one tool. See Your First Tool if you need to get started.

What You’ll Build

By the end of this guide, your server will have:
  • Per-user rate limiting on tools
  • Concurrency control to prevent resource exhaustion
  • Execution timeouts to catch hanging requests
  • IP filtering for production security
  • Redis-backed distributed rate limiting

Step 1: Add Rate Limiting to a Tool

1

Configure rate limiting on a tool

Add a rateLimit option to your tool decorator:
This limits each user to 30 search requests per minute.
2

Register the tool and enable guard

Enable the guard system in your app configuration:
A tool’s own rateLimit and concurrency are enforced without throttle. Setting throttle.enabled: true also turns on the app-level options (global, defaultRateLimit, ipFilter, …), and throttle.enabled: false turns every guard off.
3

Test the rate limit

Start your server and send rapid requests. After 30 requests within a minute, the server returns a 429 error:

Step 2: Add Concurrency Control

Prevent expensive tools from running too many instances simultaneously.
1

Add concurrency to a tool

This allows at most 2 report generations at once. Additional requests wait up to 15 seconds for a slot.
2

Understand queue behavior

When all slots are occupied:
  • With queueTimeoutMs: 0 (default), the request is immediately rejected with ConcurrencyLimitError (429).
  • With queueTimeoutMs: 15_000, the request waits up to 15 seconds. If a slot opens, it proceeds. If not, it fails with QueueTimeoutError (429).
For mutex-like behavior (only one execution at a time), set maxConcurrent: 1:

Step 3: Add Execution Timeout

Protect against hanging requests by setting a maximum execution time.
1

Add timeout to a tool

If execution takes longer than 30 seconds, it throws ExecutionTimeoutError (408) and aborts this.signal, along with the tool’s pending this.fetch() requests. Pass this.signal to other cancellable calls so the work stops too.
2

Set a default timeout for all tools

Instead of adding timeout to every tool, set a default at the app level:
Tools with their own timeout override the default. Tools without timeout use the app default.

Step 4: Global Rate Limiting

Add a server-wide rate limit that applies to all requests, regardless of which tool is called.
partitionBy: 'ip' requires FRONTMCP_TRUST_PROXY when you run behind a proxy.The client IP comes from the socket peer address by default. X-Forwarded-For and X-Real-IP are set by whoever sent the request, so they are ignored unless you declare a trusted proxy with FRONTMCP_TRUST_PROXY=true. Without that, a deployment behind a load balancer sees every request as coming from the balancer and rate-limits all clients as one. With it but no proxy actually in front, a caller can forge the header and mint a fresh bucket per request.Set FRONTMCP_TRUSTED_PROXY_DEPTH to the number of proxies you run in front of the app (default 1). The client is read from that many hops back along the chain, because callers can prepend entries but only your own proxies append to it.On Cloudflare Workers, Deno and Bun the web-fetch adapter supplies the platform’s own peer address (see Where the client IP comes from), so edge callers are partitioned per client too.When no IP can be established the request falls back to the authenticated user, and to a single ip:unresolved partition when there is no user either. It never keys on the session id: mcp-session-id is caller-supplied and a request without one is given a fresh UUID, so keying on it would mint a new budget per request. The shared bucket is contended by design — bounded contention beats an unbounded budget — and declaring your proxy is what takes callers out of it.
1

Configure global limits

Global limits are checked before per-tool limits. Both must pass for a request to proceed.
2

Combine with per-tool limits

Global and per-tool limits work independently. A tool can have its own stricter limit:
Even if the global limit allows 500 requests/min per IP, this tool is limited to 5 requests/min per user.

Step 5: IP Filtering

Block malicious IPs and restrict access to known networks. The filter is the first guard stage (checkIpFilter) of every HTTP-facing flow, so it runs before the rate-limit check and before authentication, on every route the server answers: Health and readiness probes (/healthz, /readyz, /health) are not filtered, so an orchestrator can always reach them; neither is /metrics, which has its own bearer token. An ipFilter block is enforced on its own — it does not need a global rate limit configured alongside it. Because each stage is a flow stage, a plugin can hook checkIpFilter on any of these flows.
The client IP comes from the socket peer unless a trusted proxy is declared. Behind a load balancer, set FRONTMCP_TRUST_PROXY=true (and FRONTMCP_TRUSTED_PROXY_DEPTH for more than one hop) or every request will appear to come from the balancer and your lists will match the wrong address.Note that ipFilter.trustProxy and ipFilter.trustedProxyDepth are accepted by the schema but are not read: client-IP extraction happens in the SDK context layer, before guard configuration is reachable, and setting either one logs a startup warning. Use the environment variables.

Where the client IP comes from

A request whose client IP cannot be established matches neither list, so it gets defaultAction: with 'deny' it is rejected, with 'allow' it proceeds. (Before 1.8.1 such a request was always let through.)
1

Configure IP filtering

2

Understand filter precedence

The deny list is always checked first:
  1. IP on deny list → blocked (HTTP 403 — a JSON-RPC -32001 error on the MCP endpoint, the JSON body above elsewhere)
  2. IP on allow list → allowed
  3. IP on neither list, or no client IP at all → defaultAction applies ('allow' or 'deny')
With defaultAction: 'deny', only IPs explicitly on the allow list can access your server. IpBlockedError and IpNotAllowedError are exported by @frontmcp/guard for your own code; the built-in filter answers the HTTP 403 directly rather than throwing them.
3

Enable proxy trust

Behind a load balancer or reverse proxy the client IP is the proxy’s, unless you declare the proxy as trusted. This is set through the environment, not through ipFilter:
The client is read that many hops back from the end of the X-Forwarded-For chain, because callers can prepend entries but only your own proxies append to it. A chain shorter than the configured depth was not built by those proxies, so the socket peer is used instead.
ipFilter.trustProxy and ipFilter.trustedProxyDepth exist in the schema but are not read — client-IP extraction happens in the SDK context layer, before guard configuration is reachable — and setting either logs a startup warning. Use the environment variables above.
FRONTMCP_TRUST_PROXY is only as good as your network boundary. Trusting forwarded headers means trusting whoever can set them, so two things must hold:
  • Every ingress path traverses the configured proxy chain. If a caller can reach the origin directly — a public origin IP, a peered VPC, a second ingress that skips the balancer — they choose the whole X-Forwarded-For chain, and counting hops from its end just lands on an address they picked.
  • The edge strips and rebuilds the forwarded headers. The outermost proxy must discard any inbound X-Forwarded-For and X-Real-IP and write its own, so the only entries in the chain are ones your proxies appended. With depth 1 and no X-Forwarded-For at all, X-Real-IP is used — the single-hop nginx convention. It is never consulted alongside a chain, because a caller can send both.

Step 6: Production Setup with Redis

In-memory storage works for development but does not persist across restarts or share state between server instances. Use Redis for production.
1

Configure Redis storage

All rate limit counters and semaphore tickets are stored in Redis, shared across all server instances, under keys like mcp:guard:search_documents:session-…:rl:… (a trailing : on keyPrefix is dropped; before 1.8.6 the default prefix wrote mcp:guard::…, so counters briefly split between versions during a rolling deploy).throttle.storage takes the @frontmcp/utils storage shape — type picks the backend and its options go under the matching key (redis: { config } or redis: { url }). It is not the top-level redis shape: a block without type is auto-detected from REDIS_URL / REDIS_HOST and otherwise runs in memory.
2

Decide what happens when Redis is down

Rate limits fail closed. If the throttle store is unreachable at startup, the server does not start: startup rejects with GuardStorageUnavailableError (throttle.storage (redis) is unavailable: …). That is the default in production, where fallback defaults to 'error'. If Redis goes away while the server is running, a limited call is refused with the same GuardStorageUnavailableError, not an internal error: the throttle.global check answers HTTP 503 (Retry-After: 1, JSON-RPC data.code: 'GUARD_STORAGE_UNAVAILABLE'), and a per-tool limit inside tools/call answers an isError result with _meta.code: 'GUARD_STORAGE_UNAVAILABLE' (MCP carries tool-level failures in the result, over HTTP 200).To keep serving with per-instance counters instead, say so:
With fallback: 'memory' a mid-run outage switches to per-instance counters (one warning in the log) and goes back to Redis when it answers again.The top-level redis and transport.persistence behave differently: they fall back to in-memory storage and log the failure.
3

Verify distributed behavior

With Redis storage:
  • Rate limit counters are shared across instances — a user hitting different instances still sees a single limit.
  • Semaphore tickets use atomic operations — concurrency is enforced globally.
  • Pub/sub notifications make semaphore slot release detection near-instant.
For serverless environments (Vercel, AWS Lambda), use Vercel KV or Upstash:

Testing Guard Behavior

Test that your guards work correctly using the FrontMCP testing utilities.

Testing Rate Limits

Testing Concurrency Limits

Testing Timeout


Complete Example

Here is a full app with all guard features enabled: