> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentfront.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate Limiting & Guards

> Step-by-step guide to adding rate limiting, concurrency control, and IP filtering to your FrontMCP server.

This guide walks through adding production-grade traffic controls to your FrontMCP server using the Guard system.

<Info>
  **Prerequisites:** You should have a working FrontMCP server with at least one tool. See [Your First Tool](/frontmcp/guides/your-first-tool) if you need to get started.
</Info>

## What You'll Build

By the end of this guide, your server will have:

* Per-user rate limiting on tools
* Concurrency control to prevent resource exhaustion
* Execution timeouts to catch hanging requests
* IP filtering for production security
* Redis-backed distributed rate limiting

***

## Step 1: Add Rate Limiting to a Tool

<Steps>
  <Step title="Configure rate limiting on a tool">
    Add a `rateLimit` option to your tool decorator:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    import { Tool, ToolContext } from '@frontmcp/sdk';
    import { z } from '@frontmcp/sdk';

    @Tool({
      name: 'documents:search',
      description: 'Search documents',
      inputSchema: { query: z.string(), limit: z.number().default(10) },
      rateLimit: {
        maxRequests: 30,
        windowMs: 60_000,
        partitionBy: 'userId',
      },
    })
    class SearchDocumentsTool extends ToolContext {
      async execute({ query, limit }: { query: string; limit: number }) {
        return { results: await this.get(SearchService).search(query, limit) };
      }
    }
    ```

    This limits each user to 30 search requests per minute.
  </Step>

  <Step title="Register the tool and enable guard">
    Enable the guard system in your app configuration:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    import { FrontMcp } from '@frontmcp/sdk';

    @FrontMcp({
      info: { name: 'my-server', version: '1.0.0' },
      throttle: { enabled: true },
      tools: [SearchDocumentsTool],
    })
    class MyApp {}
    ```

    <Tip>
      A tool's own `rateLimit` and `concurrency` are enforced without `throttle`. Setting `throttle.enabled: true` also turns on the app-level options (`global`, `defaultRateLimit`, `ipFilter`, …), and `throttle.enabled: false` turns every guard off.
    </Tip>
  </Step>

  <Step title="Test the rate limit">
    Start your server and send rapid requests. After 30 requests within a minute, the server returns a `429` error:

    ```json theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    {
      "code": -32000,
      "message": "Rate limit exceeded. Retry after 12 seconds."
    }
    ```
  </Step>
</Steps>

***

## Step 2: Add Concurrency Control

Prevent expensive tools from running too many instances simultaneously.

<Steps>
  <Step title="Add concurrency to a tool">
    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @Tool({
      name: 'reports:generate',
      description: 'Generate a PDF report',
      inputSchema: { reportId: z.string() },
      concurrency: {
        maxConcurrent: 2,
        queueTimeoutMs: 15_000,
      },
    })
    class GenerateReportTool extends ToolContext {
      async execute({ reportId }: { reportId: string }) {
        return await this.get(ReportService).generatePdf(reportId);
      }
    }
    ```

    This allows at most 2 report generations at once. Additional requests wait up to 15 seconds for a slot.
  </Step>

  <Step title="Understand queue behavior">
    When all slots are occupied:

    * With `queueTimeoutMs: 0` (default), the request is immediately rejected with `ConcurrencyLimitError` (429).
    * With `queueTimeoutMs: 15_000`, the request waits up to 15 seconds. If a slot opens, it proceeds. If not, it fails with `QueueTimeoutError` (429).

    For mutex-like behavior (only one execution at a time), set `maxConcurrent: 1`:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    concurrency: { maxConcurrent: 1 }
    ```
  </Step>
</Steps>

***

## Step 3: Add Execution Timeout

Protect against hanging requests by setting a maximum execution time.

<Steps>
  <Step title="Add timeout to a tool">
    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @Tool({
      name: 'llm:analyze',
      description: 'Analyze text with LLM',
      inputSchema: { text: z.string() },
      timeout: { executeMs: 30_000 },
    })
    class AnalyzeTool extends ToolContext {
      async execute({ text }: { text: string }) {
        return await this.get(LlmService).analyze(text);
      }
    }
    ```

    If execution takes longer than 30 seconds, it throws `ExecutionTimeoutError` (408) and aborts `this.signal`, along with the tool's pending `this.fetch()` requests. Pass `this.signal` to other cancellable calls so the work stops too.
  </Step>

  <Step title="Set a default timeout for all tools">
    Instead of adding `timeout` to every tool, set a default at the app level:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @FrontMcp({
      info: { name: 'my-server', version: '1.0.0' },
      throttle: {
        enabled: true,
        defaultTimeout: { executeMs: 15_000 },
      },
      tools: [AnalyzeTool, SearchDocumentsTool, GenerateReportTool],
    })
    class MyApp {}
    ```

    Tools with their own `timeout` override the default. Tools without `timeout` use the app default.
  </Step>
</Steps>

***

## Step 4: Global Rate Limiting

Add a server-wide rate limit that applies to all requests, regardless of which tool is called.

<Warning>
  **`partitionBy: 'ip'` requires `FRONTMCP_TRUST_PROXY` when you run behind a proxy.**

  The client IP comes from the socket peer address by default. `X-Forwarded-For` and
  `X-Real-IP` are set by whoever sent the request, so they are ignored unless you declare a
  trusted proxy with `FRONTMCP_TRUST_PROXY=true`. Without that, a deployment behind a load
  balancer sees every request as coming from the balancer and rate-limits all clients as one.
  With it but no proxy actually in front, a caller can forge the header and mint a fresh
  bucket per request.

  Set `FRONTMCP_TRUSTED_PROXY_DEPTH` to the number of proxies you run in front of the app
  (default `1`). The client is read from that many hops back along the chain, because callers
  can prepend entries but only your own proxies append to it.

  On Cloudflare Workers, Deno and Bun the web-fetch adapter supplies the platform's own peer
  address (see [Where the client IP comes from](#where-the-client-ip-comes-from)), so edge
  callers are partitioned per client too.

  When no IP can be established the request falls back to the authenticated user, and to a
  single `ip:unresolved` partition when there is no user either. It never keys on the session
  id: `mcp-session-id` is caller-supplied and a request without one is given a fresh UUID, so
  keying on it would mint a new budget per request. The shared bucket is contended by design --
  bounded contention beats an unbounded budget -- and declaring your proxy is what takes
  callers out of it.
</Warning>

<Steps>
  <Step title="Configure global limits">
    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @FrontMcp({
      info: { name: 'my-server', version: '1.0.0' },
      throttle: {
        enabled: true,
        global: {
          maxRequests: 500,
          windowMs: 60_000,
          partitionBy: 'ip',
        },
        globalConcurrency: {
          maxConcurrent: 20,
        },
      },
      tools: [AnalyzeTool, SearchDocumentsTool, GenerateReportTool],
    })
    class MyApp {}
    ```

    Global limits are checked **before** per-tool limits. Both must pass for a request to proceed.
  </Step>

  <Step title="Combine with per-tool limits">
    Global and per-tool limits work independently. A tool can have its own stricter limit:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @Tool({
      name: 'expensive:operation',
      inputSchema: { id: z.string() },
      rateLimit: { maxRequests: 5, windowMs: 60_000, partitionBy: 'userId' },
    })
    class ExpensiveTool extends ToolContext { /* ... */ }
    ```

    Even if the global limit allows 500 requests/min per IP, this tool is limited to 5 requests/min per user.
  </Step>
</Steps>

***

## Step 5: IP Filtering

Block malicious IPs and restrict access to known networks.

The filter is the first guard stage (`checkIpFilter`) of every HTTP-facing flow, so it runs
before the rate-limit check and before authentication, on every route the server answers:

| Route | Rejection |
| - | - |
| MCP endpoint (`<entryPath>`, `/sse`, `/message`) | HTTP 403, JSON-RPC error `-32001` |
| `/oauth/*` (authorize, token, register, callback, provider callback, connect, userinfo, UI) | HTTP 403, JSON `{ "error": "forbidden", "message": "Client IP rejected by ipFilter" }` |
| `/.well-known/*` (protected resource, authorization server, JWKS) | same JSON 403 |
| `llm.txt`, `llm_full.txt` and the skills HTTP API | same JSON 403 |
| Custom `http.routes` (checked before `auth: true` verification) | same JSON 403 |

Health and readiness probes (`/healthz`, `/readyz`, `/health`) are **not** filtered, so an
orchestrator can always reach them; neither is `/metrics`, which has its own bearer token. An
`ipFilter` block is enforced on its own -- it does not need a `global` rate limit configured
alongside it. Because each stage is a flow stage, a plugin can hook `checkIpFilter` on any of
these flows.

<Warning>
  The client IP comes from the socket peer unless a trusted proxy is declared. Behind a load
  balancer, set `FRONTMCP_TRUST_PROXY=true` (and `FRONTMCP_TRUSTED_PROXY_DEPTH` for more than
  one hop) or every request will appear to come from the balancer and your lists will match
  the wrong address.

  Note that `ipFilter.trustProxy` and `ipFilter.trustedProxyDepth` are accepted by the schema
  but are **not** read: client-IP extraction happens in the SDK context layer, before guard
  configuration is reachable, and setting either one logs a startup warning. Use the
  environment variables.
</Warning>

### Where the client IP comes from

| Runtime | Client IP |
| - | - |
| Node (Express, serverless handler) | The socket peer. A dual-stack socket reports IPv4 clients as `::ffff:a.b.c.d`; the filter matches that against IPv4 rules. |
| Behind a proxy you declare | `X-Forwarded-For`, counted `FRONTMCP_TRUSTED_PROXY_DEPTH` hops from the end, when `FRONTMCP_TRUST_PROXY=true`. |
| Cloudflare Workers (incl. Durable Objects) | `CF-Connecting-IP`, which Cloudflare's edge sets on every request. Trusted only when the runtime is Workers; on any other runtime the header is ignored. |
| Deno | `info.remoteAddr.hostname` — pass the handler to `Deno.serve(handler)` or forward `info` as its second argument. |
| Bun | `server.requestIP(request)` — use `Bun.serve({ fetch: handler })` or forward `server` as its second argument. |

A request whose client IP cannot be established matches neither list, so it gets
`defaultAction`: with `'deny'` it is rejected, with `'allow'` it proceeds. (Before 1.8.1 such a
request was always let through.)

<Steps>
  <Step title="Configure IP filtering">
    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @FrontMcp({
      info: { name: 'my-server', version: '1.0.0' },
      throttle: {
        enabled: true,
        ipFilter: {
          denyList: [
            '192.0.2.1',            // Known bad actor
            '198.51.100.0/24',      // Blocked subnet
          ],
          allowList: [
            '10.0.0.0/8',           // Internal network
            '172.16.0.0/12',        // Office VPN
            '2001:db8::/32',        // IPv6 office range
          ],
          defaultAction: 'deny',    // Block everything not on allowList
        },
      },
      tools: [MyTool],
    })
    class MyApp {}
    ```
  </Step>

  <Step title="Understand filter precedence">
    The deny list is always checked first:

    1. IP on deny list → **blocked** (HTTP 403 -- a JSON-RPC `-32001` error on the MCP endpoint, the JSON body above elsewhere)
    2. IP on allow list → **allowed**
    3. IP on neither list, or no client IP at all → `defaultAction` applies (`'allow'` or `'deny'`)

    With `defaultAction: 'deny'`, only IPs explicitly on the allow list can access your server.
    `IpBlockedError` and `IpNotAllowedError` are exported by `@frontmcp/guard` for your own code;
    the built-in filter answers the HTTP 403 directly rather than throwing them.
  </Step>

  <Step title="Enable proxy trust">
    Behind a load balancer or reverse proxy the client IP is the proxy's, unless you declare
    the proxy as trusted. This is set through the environment, not through `ipFilter`:

    ```bash theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    FRONTMCP_TRUST_PROXY=true
    FRONTMCP_TRUSTED_PROXY_DEPTH=2   # behind 2 proxies (e.g. CloudFront + ALB)
    ```

    The client is read that many hops back from the end of the `X-Forwarded-For` chain,
    because callers can prepend entries but only your own proxies append to it. A chain
    shorter than the configured depth was not built by those proxies, so the socket peer is
    used instead.

    <Note>
      `ipFilter.trustProxy` and `ipFilter.trustedProxyDepth` exist in the schema but are not
      read — client-IP extraction happens in the SDK context layer, before guard
      configuration is reachable — and setting either logs a startup warning. Use the
      environment variables above.
    </Note>

    <Warning>
      **`FRONTMCP_TRUST_PROXY` is only as good as your network boundary.** Trusting forwarded
      headers means trusting whoever can set them, so two things must hold:

      * **Every ingress path traverses the configured proxy chain.** If a caller can reach the
        origin directly -- a public origin IP, a peered VPC, a second ingress that skips the
        balancer -- they choose the whole `X-Forwarded-For` chain, and counting hops from its end
        just lands on an address they picked.
      * **The edge strips and rebuilds the forwarded headers.** The outermost proxy must discard
        any inbound `X-Forwarded-For` and `X-Real-IP` and write its own, so the only entries in the
        chain are ones your proxies appended.

        With depth `1` and no `X-Forwarded-For` at all, `X-Real-IP` is used -- the single-hop nginx
        convention. It is never consulted alongside a chain, because a caller can send both.
    </Warning>
  </Step>
</Steps>

***

## Step 6: Production Setup with Redis

In-memory storage works for development but does not persist across restarts or share state between server instances. Use Redis for production.

<Steps>
  <Step title="Configure Redis storage">
    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    @FrontMcp({
      info: { name: 'production-server', version: '1.0.0' },
      throttle: {
        enabled: true,
        storage: {
          type: 'redis',
          redis: {
            config: {
              host: process.env.REDIS_HOST ?? 'localhost',
              port: Number(process.env.REDIS_PORT ?? 6379),
              password: process.env.REDIS_PASSWORD,
              tls: process.env.NODE_ENV === 'production',
            },
          },
        },
        keyPrefix: 'mcp:guard:',
        global: { maxRequests: 1000, windowMs: 60_000, partitionBy: 'ip' },
        defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: 'session' },
        defaultConcurrency: { maxConcurrent: 10 },
        defaultTimeout: { executeMs: 30_000 },
      },
      tools: [SearchDocumentsTool, GenerateReportTool, AnalyzeTool],
    })
    class ProductionApp {}
    ```

    All rate limit counters and semaphore tickets are stored in Redis, shared across all server instances, under keys like `mcp:guard:search_documents:session-…:rl:…` (a trailing `:` on `keyPrefix` is dropped; before 1.8.6 the default prefix wrote `mcp:guard::…`, so counters briefly split between versions during a rolling deploy).

    `throttle.storage` takes the `@frontmcp/utils` storage shape — `type` picks the backend and its options go under the matching key (`redis: { config }` or `redis: { url }`). It is **not** the top-level `redis` shape: a block without `type` is auto-detected from `REDIS_URL` / `REDIS_HOST` and otherwise runs in memory.
  </Step>

  <Step title="Decide what happens when Redis is down">
    Rate limits fail closed. If the throttle store is unreachable at startup, the server does not start: startup rejects with `GuardStorageUnavailableError` (`throttle.storage (redis) is unavailable: …`). That is the default in production, where `fallback` defaults to `'error'`. If Redis goes away while the server is running, a limited call is refused with the same `GuardStorageUnavailableError` (a readable 503), not an internal error.

    To keep serving with per-instance counters instead, say so:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    storage: {
      type: 'redis',
      redis: { url: process.env.REDIS_URL },
      fallback: 'memory',
    },
    ```

    With `fallback: 'memory'` a mid-run outage switches to per-instance counters (one warning in the log) and goes back to Redis when it answers again.

    The top-level `redis` and `transport.persistence` behave differently: they fall back to in-memory storage and log the failure.
  </Step>

  <Step title="Verify distributed behavior">
    With Redis storage:

    * Rate limit counters are shared across instances — a user hitting different instances still sees a single limit.
    * Semaphore tickets use atomic operations — concurrency is enforced globally.
    * Pub/sub notifications make semaphore slot release detection near-instant.

    For serverless environments (Vercel, AWS Lambda), use Vercel KV or Upstash:

    ```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
    storage: {
      type: 'vercel-kv',
      vercelKv: {
        url: process.env.KV_REST_API_URL,
        token: process.env.KV_REST_API_TOKEN,
      },
    },
    ```
  </Step>
</Steps>

***

## Testing Guard Behavior

Test that your guards work correctly using the FrontMCP testing utilities.

### Testing Rate Limits

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
import { connect } from '@frontmcp/sdk';
import { MyApp } from './app';

describe('SearchDocumentsTool rate limiting', () => {
  it('should reject after exceeding rate limit', async () => {
    const client = await connect(MyApp);

    // Send requests up to the limit
    for (let i = 0; i < 30; i++) {
      const result = await client.callTool('documents:search', { query: 'test' });
      expect(result.isError).toBe(false);
    }

    // Next request should be rate-limited
    const result = await client.callTool('documents:search', { query: 'test' });
    expect(result.isError).toBe(true);
  });
});
```

### Testing Concurrency Limits

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
describe('GenerateReportTool concurrency', () => {
  it('should limit concurrent executions', async () => {
    const client = await connect(MyApp);

    // Start 3 concurrent requests (limit is 2, no queue)
    const results = await Promise.allSettled([
      client.callTool('reports:generate', { reportId: '1' }),
      client.callTool('reports:generate', { reportId: '2' }),
      client.callTool('reports:generate', { reportId: '3' }),
    ]);

    const rejected = results.filter((r) => r.status === 'rejected');
    expect(rejected.length).toBeGreaterThanOrEqual(1);
  });
});
```

### Testing Timeout

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
describe('AnalyzeTool timeout', () => {
  it('should timeout on slow execution', async () => {
    // Mock a slow service
    jest.spyOn(LlmService.prototype, 'analyze').mockImplementation(
      () => new Promise((resolve) => setTimeout(resolve, 60_000)),
    );

    const client = await connect(MyApp);
    const result = await client.callTool('llm:analyze', { text: 'test' });
    expect(result.isError).toBe(true);
  });
});
```

***

## Complete Example

Here is a full app with all guard features enabled:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
import { FrontMcp, Tool, ToolContext } from '@frontmcp/sdk';
import { z } from '@frontmcp/sdk';

@Tool({
  name: 'search',
  description: 'Search documents',
  inputSchema: { query: z.string() },
  rateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: 'userId' },
  timeout: { executeMs: 10_000 },
})
class SearchTool extends ToolContext {
  async execute({ query }: { query: string }) {
    return { results: [] };
  }
}

@Tool({
  name: 'generate-report',
  description: 'Generate PDF report',
  inputSchema: { id: z.string() },
  rateLimit: { maxRequests: 10, windowMs: 60_000, partitionBy: 'userId' },
  concurrency: { maxConcurrent: 2, queueTimeoutMs: 10_000 },
  timeout: { executeMs: 60_000 },
})
class ReportTool extends ToolContext {
  async execute({ id }: { id: string }) {
    return { url: `/reports/${id}.pdf` };
  }
}

@FrontMcp({
  info: { name: 'guarded-server', version: '1.0.0' },
  throttle: {
    enabled: true,
    storage: {
      type: 'redis',
      redis: { config: { host: process.env.REDIS_HOST ?? 'localhost', port: 6379 } },
    },
    global: { maxRequests: 1000, windowMs: 60_000, partitionBy: 'ip' },
    defaultTimeout: { executeMs: 30_000 },
    ipFilter: {
      denyList: ['192.0.2.0/24'],
    },
  },
  tools: [SearchTool, ReportTool],
})
class GuardedServer {}
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.