Prerequisites: You should have a working FrontMCP server with at least one tool. See Your First Tool if you need to get started.
What You’ll Build
By the end of this guide, your server will have:- Per-user rate limiting on tools
- Concurrency control to prevent resource exhaustion
- Execution timeouts to catch hanging requests
- IP filtering for production security
- Redis-backed distributed rate limiting
Step 1: Add Rate Limiting to a Tool
1
Configure rate limiting on a tool
Add a This limits each user to 30 search requests per minute.
rateLimit option to your tool decorator:2
Register the tool and enable guard
Enable the guard system in your app configuration:
3
Test the rate limit
Start your server and send rapid requests. After 30 requests within a minute, the server returns a
429 error:Step 2: Add Concurrency Control
Prevent expensive tools from running too many instances simultaneously.1
Add concurrency to a tool
2
Understand queue behavior
When all slots are occupied:
- With
queueTimeoutMs: 0(default), the request is immediately rejected withConcurrencyLimitError(429). - With
queueTimeoutMs: 15_000, the request waits up to 15 seconds. If a slot opens, it proceeds. If not, it fails withQueueTimeoutError(429).
maxConcurrent: 1:Step 3: Add Execution Timeout
Protect against hanging requests by setting a maximum execution time.1
Add timeout to a tool
ExecutionTimeoutError (408) and aborts this.signal, along with the tool’s pending this.fetch() requests. Pass this.signal to other cancellable calls so the work stops too.2
Set a default timeout for all tools
Instead of adding Tools with their own
timeout to every tool, set a default at the app level:timeout override the default. Tools without timeout use the app default.Step 4: Global Rate Limiting
Add a server-wide rate limit that applies to all requests, regardless of which tool is called.1
Configure global limits
2
Combine with per-tool limits
Global and per-tool limits work independently. A tool can have its own stricter limit:Even if the global limit allows 500 requests/min per IP, this tool is limited to 5 requests/min per user.
Step 5: IP Filtering
Block malicious IPs and restrict access to known networks. The filter is the first guard stage (checkIpFilter) of every HTTP-facing flow, so it runs
before the rate-limit check and before authentication, on every route the server answers:
Health and readiness probes (
/healthz, /readyz, /health) are not filtered, so an
orchestrator can always reach them; neither is /metrics, which has its own bearer token. An
ipFilter block is enforced on its own — it does not need a global rate limit configured
alongside it. Because each stage is a flow stage, a plugin can hook checkIpFilter on any of
these flows.
Where the client IP comes from
A request whose client IP cannot be established matches neither list, so it gets
defaultAction: with 'deny' it is rejected, with 'allow' it proceeds. (Before 1.8.1 such a
request was always let through.)
1
Configure IP filtering
2
Understand filter precedence
The deny list is always checked first:
- IP on deny list → blocked (HTTP 403 — a JSON-RPC
-32001error on the MCP endpoint, the JSON body above elsewhere) - IP on allow list → allowed
- IP on neither list, or no client IP at all →
defaultActionapplies ('allow'or'deny')
defaultAction: 'deny', only IPs explicitly on the allow list can access your server.
IpBlockedError and IpNotAllowedError are exported by @frontmcp/guard for your own code;
the built-in filter answers the HTTP 403 directly rather than throwing them.3
Enable proxy trust
Behind a load balancer or reverse proxy the client IP is the proxy’s, unless you declare
the proxy as trusted. This is set through the environment, not through The client is read that many hops back from the end of the
ipFilter:X-Forwarded-For chain,
because callers can prepend entries but only your own proxies append to it. A chain
shorter than the configured depth was not built by those proxies, so the socket peer is
used instead.ipFilter.trustProxy and ipFilter.trustedProxyDepth exist in the schema but are not
read — client-IP extraction happens in the SDK context layer, before guard
configuration is reachable — and setting either logs a startup warning. Use the
environment variables above.Step 6: Production Setup with Redis
In-memory storage works for development but does not persist across restarts or share state between server instances. Use Redis for production.1
Configure Redis storage
mcp:guard:search_documents:session-…:rl:… (a trailing : on keyPrefix is dropped; before 1.8.6 the default prefix wrote mcp:guard::…, so counters briefly split between versions during a rolling deploy).throttle.storage takes the @frontmcp/utils storage shape — type picks the backend and its options go under the matching key (redis: { config } or redis: { url }). It is not the top-level redis shape: a block without type is auto-detected from REDIS_URL / REDIS_HOST and otherwise runs in memory.2
Decide what happens when Redis is down
Rate limits fail closed. If the throttle store is unreachable at startup, the server does not start: startup rejects with With
GuardStorageUnavailableError (throttle.storage (redis) is unavailable: …). That is the default in production, where fallback defaults to 'error'. If Redis goes away while the server is running, a limited call is refused with the same GuardStorageUnavailableError (a readable 503), not an internal error.To keep serving with per-instance counters instead, say so:fallback: 'memory' a mid-run outage switches to per-instance counters (one warning in the log) and goes back to Redis when it answers again.The top-level redis and transport.persistence behave differently: they fall back to in-memory storage and log the failure.3
Verify distributed behavior
With Redis storage:
- Rate limit counters are shared across instances — a user hitting different instances still sees a single limit.
- Semaphore tickets use atomic operations — concurrency is enforced globally.
- Pub/sub notifications make semaphore slot release detection near-instant.