Skip to main content
Guard provides rate limiting, concurrency control, execution timeout, and IP filtering for your MCP server. It protects against abuse, ensures fair resource allocation, and prevents runaway requests.
Guard is powered by the @frontmcp/guard library and integrates directly into tool and agent flows. All guard checks run automatically before execution, with cleanup handled in finalize stages.

Why Guard?


Quick Start

Add rate limiting and a timeout to any tool with decorator options:
A tool’s own rateLimit and concurrency are enforced without a throttle option, as are those of an agent, of the tools declared inside an @Agent and of its nested agents. Set throttle.enabled: false to turn every guard off, including these.

How Guard Integrates with Flows

Guard checks are implemented as flow stages that run automatically in the tool and agent execution pipelines:
  1. acquireQuota — Checks global and per-entity rate limits. Throws RateLimitError if exceeded. Over HTTP, throttle.global is checked once per request by the http:request flow; stdio and direct calls check it here.
  2. acquireSemaphore — Takes a slot from throttle.globalConcurrency and from the entity’s concurrency (else throttle.defaultConcurrency). Throws ConcurrencyLimitError if no slot is available. A tool called from another tool with this.callTool(), and a tool or agent an agent calls during its run, runs inside its caller’s global slot and takes only its own concurrency slot.
  3. execute — Wrapped with withTimeout if a timeout is configured. Throws ExecutionTimeoutError if exceeded.
  4. releaseSemaphore — Releases the concurrency slot back to the pool. After a timeout, the slot stays taken until the timed-out execute() actually returns.
  5. releaseQuota — Cleans up rate limit state.

Rate Limiting

FrontMCP uses a sliding window algorithm for rate limiting. It provides smooth, accurate throttling with O(1) storage per key.

Per-Tool Rate Limiting

Global Rate Limiting

Set a server-wide rate limit in your app configuration:
The global rate limit is checked before per-entity limits. Both must pass for a request to proceed.

Partition Strategies

Partition keys determine how rate limits are bucketed: A 'session' bucket is only ever a session the server verified: a mcp-session-id the server does not accept is ignored. A request without a verified session (every MCP 2026-07-28 request, stateless HTTP, or a session id the server rejected) is not given a fresh bucket: it falls back to the signed-in user, and anonymous callers share one anonymous bucket. 'userId' never counts an anonymous caller as a user; it falls back to the session the same way. throttle.global partitioned by 'ip' or 'global' is checked before authentication. Partitioned by 'session', 'userId' or a function, it is checked right after authentication, on the verified identity. Custom partition key example:

Concurrency Control

Concurrency control uses a distributed semaphore to limit how many instances of a tool or agent can execute simultaneously.

Per-Tool Concurrency

Mutex Pattern

Set maxConcurrent: 1 to ensure only one execution at a time:

Queue Behavior

When queueTimeoutMs is set, requests that cannot acquire a slot immediately will wait in a queue:
  • queueTimeoutMs: 0 (default) — Immediately reject if no slot available. Throws ConcurrencyLimitError.
  • queueTimeoutMs: 5000 — Wait up to 5 seconds for a slot. Throws QueueTimeoutError if the wait expires.
The semaphore uses pub/sub notifications when available (Redis) for efficient slot release detection, falling back to polling with exponential backoff.

Execution Timeout

Timeout wraps the execute stage with a deadline. If execution exceeds the configured duration, it throws ExecutionTimeoutError. When the deadline passes, a tool’s this.signal is aborted, and so are the requests the tool made with this.fetch(). Pass the signal to other cancellable work too: the deadline answers the client, but it cannot stop code that ignores the signal. Until execute() returns, it keeps its concurrency slot, so a maxConcurrent limit still holds.

Per-Tool Timeout

Default Timeout

Set a default timeout for all tools and agents at the app level:
Per-entity timeout takes precedence over the app default.

IP Filtering

IP filtering allows or blocks requests based on client IP address, supporting IPv4, IPv6, and CIDR ranges.

Filter Precedence

  1. Deny list is checked first. If matched, the request is blocked with HTTP 403.
  2. Allow list is checked next. If matched, the request proceeds.
  3. Default action applies if neither list matches, and when no client IP could be established:
    • 'allow' (default) — Request proceeds.
    • 'deny' — Request is blocked with HTTP 403.
A blocked MCP request (<entryPath>, /sse, /message) gets JSON-RPC error -32001. Every other route answers { "error": "forbidden", "message": "Client IP rejected by ipFilter" }.

Route Coverage

The filter is the checkIpFilter stage at the start of every HTTP-facing flow, ahead of rate limiting and authentication: the MCP endpoint, /oauth/*, /.well-known/*, llm.txt / llm_full.txt, the skills HTTP API, and custom http.routes (through the http:ip-filter flow, before auth: true verification). Health and readiness probes (/healthz, /readyz, /health) and /metrics (bearer-token protected) are exempt.

Client IP Sources

Supported IP Formats

Proxy Configuration

When your server is behind a reverse proxy (Nginx, CloudFront, etc.), declare the proxy as trusted so the client IP is read from X-Forwarded-For instead of the socket peer. This is set through the environment:
The client is read that many hops back from the end of the chain, because callers can prepend entries but only your own proxies append to it. A chain shorter than the configured depth is ignored in favour of the socket peer.
ipFilter.trustProxy and ipFilter.trustedProxyDepth exist in the schema but are not read: client-IP extraction happens in the SDK context layer, before guard configuration is reachable, and setting either logs a startup warning. Use the environment variables above.
FRONTMCP_TRUST_PROXY is only as good as your network boundary. Trusting forwarded headers means trusting whoever can set them, so two things must hold:
  • Every ingress path traverses the configured proxy chain. If a caller can reach the origin directly — a public origin IP, a peered VPC, a second ingress that skips the balancer — they choose the whole X-Forwarded-For chain, and counting hops from its end just lands on an address they picked.
  • The edge strips and rebuilds the forwarded headers. The outermost proxy must discard any inbound X-Forwarded-For and X-Real-IP and write its own, so the only entries in the chain are ones your proxies appended. With depth 1 and no X-Forwarded-For at all, X-Real-IP is used — the single-hop nginx convention. It is never consulted alongside a chain, because a caller can send both.

App-Level Configuration

The throttle field in @FrontMcp configures all guard features at the app level:

Configuration Precedence


Storage Backends

Guard supports multiple storage backends for distributed deployments.

Memory (Development)

The default backend. Suitable for single-instance development. No configuration needed.
In-memory storage does not persist across restarts and does not work with multiple server instances. Use Redis for production.

Redis (Production)

For distributed rate limiting across multiple server instances:
storage is a StorageConfig from @frontmcp/utils, not the top-level redis shape: the backend is chosen by type, and its connection goes under the key of that name (redis: { config } or redis: { url: process.env.REDIS_URL }). A block without type (such as { provider: 'redis', host, port }) is auto-detected from REDIS_URL / REDIS_HOST and otherwise runs in memory. Redis enables pub/sub-based semaphore notifications for more efficient concurrency slot release detection.

Vercel KV / Upstash

For serverless environments:
Upstash is { type: 'upstash', upstash: { url, token } }.

When the store is unreachable

Rate limits are a security control, so they fail closed. If the configured backend cannot be reached at startup, the server does not start: createGuardManager rejects with GuardStorageUnavailableError (GUARD_STORAGE_UNAVAILABLE), whose message names throttle.storage. This is the default in production, where fallback defaults to 'error' (it defaults to 'memory' otherwise). The same applies if the backend goes away while the server is running: a rate-limited or concurrency-limited call is refused with the same GuardStorageUnavailableError (503, GUARD_STORAGE_UNAVAILABLE, a public error the client can read), not an internal error. What the client receives depends on where the check runs. The throttle.global limit is checked on the HTTP request, so it answers HTTP 503 with Retry-After: 1 and a JSON-RPC error whose data.code is GUARD_STORAGE_UNAVAILABLE. A per-tool rateLimit / concurrency check runs inside tools/call, and MCP returns tool-level failures in the result: HTTP 200 with isError: true and _meta.code: 'GUARD_STORAGE_UNAVAILABLE'. Either way the client reads “Service temporarily unavailable: the rate-limit store cannot be reached”; the store’s address and the reason stay in the server log. The Redis client logs the connection error at a rate-limited interval instead of once per reconnect attempt. To keep serving with per-instance counters while the store is down, opt in explicitly:
This differs from the top-level redis and transport.persistence, which fall back to in-memory storage with an error log.

Key format

Keys are <keyPrefix><entity>:<partition>:<kind>:…, for example mcp:guard:export_tickets:global:rl:1790722980000. A trailing : on keyPrefix is dropped, because the storage namespace adds its own separator.
Before 1.8.6 the default prefix produced a double colon (mcp:guard::export_tickets:…). Counters written by an older instance are not read by a newer one, so during a rolling deploy to 1.8.6 each version counts separately until the old instances are gone, and a window’s limit can briefly be exceeded. A custom keyPrefix without a trailing colon keeps its keys.

Error Handling

Guard throws specific error classes when limits are exceeded: A tools/call that hits a limit answers with an error result whose _meta.code is the code above, in production too. The built-in ipFilter does not throw: it answers the HTTP request with 403 before any tool runs (see Filter Precedence).

Agent Guard

Agents support the same guard options as tools:
The agent flow follows the same stage ordering: acquireQuota → acquireSemaphore → execute (with timeout) → releaseSemaphore → releaseQuota.

Configuration Reference

RateLimitConfig

ConcurrencyConfig

TimeoutConfig

IpFilterConfig

GuardConfig (App-Level)