> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentfront.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Guard

> Rate limiting, concurrency control, execution timeout, and IP filtering for FrontMCP tools and agents.

Guard provides **rate limiting**, **concurrency control**, **execution timeout**, and **IP filtering** for your MCP server. It protects against abuse, ensures fair resource allocation, and prevents runaway requests.

<Info>
  Guard is powered by the `@frontmcp/guard` library and integrates directly into tool and agent flows. All guard checks run automatically before execution, with cleanup handled in finalize stages.
</Info>

## Why Guard?

| Threat | Without Guard | With Guard |
| - | - | - |
| **Client flooding requests** | Server overwhelmed | Rate-limited per user/IP |
| **Tool running forever** | Hangs, resource leak | Timeout protection |
| **Unbounded parallelism** | Resource exhaustion | Controlled concurrency |
| **Malicious IPs** | Open access | IP allow/deny filtering |

***

## Quick Start

Add rate limiting and a timeout to any tool with decorator options:

<CodeGroup>
  ```typescript Class Style theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
  import { Tool, ToolContext } from '@frontmcp/sdk';
  import { z } from '@frontmcp/sdk';

  @Tool({
  name: 'search',
  description: 'Search documents',
  inputSchema: { query: z.string() },
  rateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: 'userId' },
  timeout: { executeMs: 10_000 },
  })
  class SearchTool extends ToolContext {
  async execute({ query }: { query: string }) {
  return { results: await this.get(SearchService).search(query) };
  }
  }

  ```

  ```typescript Function Style theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
  import { tool } from '@frontmcp/sdk';
  import { z } from '@frontmcp/sdk';

  const SearchTool = tool({
    name: 'search',
    description: 'Search documents',
    inputSchema: { query: z.string() },
    rateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: 'userId' },
    timeout: { executeMs: 10_000 },
  })(async ({ query }, ctx) => {
    return { results: await ctx.get(SearchService).search(query) };
  });
  ```
</CodeGroup>

A tool's own `rateLimit` and `concurrency` are enforced without a `throttle` option. Set `throttle.enabled: false` to turn every guard off, including these.

***

## How Guard Integrates with Flows

Guard checks are implemented as flow stages that run automatically in the tool and agent execution pipelines:

```
Pre stages:      ... → acquireQuota → acquireSemaphore → ...
Execute stages:  validateInput → execute (wrapped with timeout) → validateOutput
Finalize stages: releaseSemaphore → releaseQuota → ...
```

1. **acquireQuota** — Checks global and per-entity rate limits. Throws `RateLimitError` if exceeded. Over HTTP, `throttle.global` is checked once per request by the `http:request` flow; stdio and direct calls check it here.
2. **acquireSemaphore** — Takes a slot from `throttle.globalConcurrency` and from the entity's `concurrency` (else `throttle.defaultConcurrency`). Throws `ConcurrencyLimitError` if no slot is available. A tool called from another tool with `this.callTool()` runs inside its caller's global slot and takes only its own `concurrency` slot.
3. **execute** — Wrapped with `withTimeout` if a timeout is configured. Throws `ExecutionTimeoutError` if exceeded.
4. **releaseSemaphore** — Releases the concurrency slot back to the pool. After a timeout, the slot stays taken until the timed-out `execute()` actually returns.
5. **releaseQuota** — Cleans up rate limit state.

***

## Rate Limiting

FrontMCP uses a **sliding window** algorithm for rate limiting. It provides smooth, accurate throttling with O(1) storage per key.

### Per-Tool Rate Limiting

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Tool({
  name: 'api:call',
  inputSchema: { endpoint: z.string() },
  rateLimit: {
    maxRequests: 100,       // 100 requests
    windowMs: 60_000,       // per 60 seconds
    partitionBy: 'userId',  // per user
  },
})
class ApiCallTool extends ToolContext {
  async execute({ endpoint }: { endpoint: string }) {
    return await this.get(ApiService).call(endpoint);
  }
}
```

### Global Rate Limiting

Set a server-wide rate limit in your app configuration:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@FrontMcp({
  name: 'my-server',
  throttle: {
    enabled: true,
    global: {
      maxRequests: 1000,
      windowMs: 60_000,
      partitionBy: 'ip',
    },
  },
  tools: [ApiCallTool, SearchTool],
})
class MyApp {}
```

The global rate limit is checked **before** per-entity limits. Both must pass for a request to proceed.

### Partition Strategies

Partition keys determine how rate limits are bucketed:

| Strategy | Description | Use Case |
| - | - | - |
| `'global'` | Single shared bucket | Server-wide limits |
| `'ip'` | Per client IP address | Prevent IP-based abuse |
| `'session'` | Per MCP session ID | Per-connection limits |
| `'userId'` | Per authenticated user | Per-user quotas |
| Custom function | `(ctx) => string` | Tenant, org, or custom grouping |

A `'session'` bucket is only ever a session the server verified: a `mcp-session-id` the server does not accept is ignored. A request without a verified session (every MCP 2026-07-28 request, stateless HTTP, or a session id the server rejected) is not given a fresh bucket: it falls back to the signed-in user, and anonymous callers share one `anonymous` bucket. `'userId'` never counts an anonymous caller as a user; it falls back to the session the same way.

`throttle.global` partitioned by `'ip'` or `'global'` is checked before authentication. Partitioned by `'session'`, `'userId'` or a function, it is checked right after authentication, on the verified identity.

**Custom partition key example:**

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Tool({
  name: 'tenant:query',
  inputSchema: { query: z.string() },
  rateLimit: {
    maxRequests: 500,
    windowMs: 60_000,
    partitionBy: (ctx) => ctx.userId?.split(':')[0] ?? 'anonymous',
  },
})
class TenantQueryTool extends ToolContext { /* ... */ }
```

***

## Concurrency Control

Concurrency control uses a **distributed semaphore** to limit how many instances of a tool or agent can execute simultaneously.

### Per-Tool Concurrency

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Tool({
  name: 'report:generate',
  inputSchema: { reportId: z.string() },
  concurrency: {
    maxConcurrent: 3,        // At most 3 simultaneous executions
    queueTimeoutMs: 10_000,  // Wait up to 10s for a slot
    partitionBy: 'global',   // Shared across all users
  },
})
class GenerateReportTool extends ToolContext {
  async execute({ reportId }: { reportId: string }) {
    return await this.get(ReportService).generate(reportId);
  }
}
```

### Mutex Pattern

Set `maxConcurrent: 1` to ensure only one execution at a time:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Tool({
  name: 'db:migrate',
  inputSchema: { version: z.string() },
  concurrency: { maxConcurrent: 1 },
})
class MigrateTool extends ToolContext { /* ... */ }
```

### Queue Behavior

When `queueTimeoutMs` is set, requests that cannot acquire a slot immediately will wait in a queue:

* **`queueTimeoutMs: 0`** (default) — Immediately reject if no slot available. Throws `ConcurrencyLimitError`.
* **`queueTimeoutMs: 5000`** — Wait up to 5 seconds for a slot. Throws `QueueTimeoutError` if the wait expires.

The semaphore uses pub/sub notifications when available (Redis) for efficient slot release detection, falling back to polling with exponential backoff.

***

## Execution Timeout

Timeout wraps the `execute` stage with a deadline. If execution exceeds the configured duration, it throws `ExecutionTimeoutError`.

When the deadline passes, a tool's `this.signal` is aborted, and so are the requests the tool made with `this.fetch()`. Pass the signal to other cancellable work too: the deadline answers the client, but it cannot stop code that ignores the signal. Until `execute()` returns, it keeps its concurrency slot, so a `maxConcurrent` limit still holds.

### Per-Tool Timeout

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Tool({
  name: 'llm:summarize',
  inputSchema: { text: z.string() },
  timeout: { executeMs: 30_000 },  // 30-second deadline
})
class SummarizeTool extends ToolContext {
  async execute({ text }: { text: string }) {
    return await this.get(LlmService).summarize(text);
  }
}
```

### Default Timeout

Set a default timeout for all tools and agents at the app level:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@FrontMcp({
  name: 'my-server',
  throttle: {
    enabled: true,
    defaultTimeout: { executeMs: 15_000 },
  },
  tools: [SummarizeTool, SearchTool],
})
class MyApp {}
```

Per-entity timeout takes precedence over the app default.

***

## IP Filtering

IP filtering allows or blocks requests based on client IP address, supporting IPv4, IPv6, and CIDR ranges.

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@FrontMcp({
  name: 'my-server',
  throttle: {
    enabled: true,
    ipFilter: {
      allowList: ['10.0.0.0/8', '172.16.0.0/12'],
      denyList: ['192.0.2.1', '198.51.100.0/24'],
      defaultAction: 'deny',
    },
  },
  tools: [MyTool],
})
class MyApp {}
```

### Filter Precedence

1. **Deny list** is checked first. If matched, the request is blocked with HTTP 403.
2. **Allow list** is checked next. If matched, the request proceeds.
3. **Default action** applies if neither list matches, and when no client IP could be established:
   * `'allow'` (default) — Request proceeds.
   * `'deny'` — Request is blocked with HTTP 403.

A blocked MCP request (`<entryPath>`, `/sse`, `/message`) gets JSON-RPC error `-32001`. Every
other route answers `{ "error": "forbidden", "message": "Client IP rejected by ipFilter" }`.

### Route Coverage

The filter is the `checkIpFilter` stage at the start of every HTTP-facing flow, ahead of rate
limiting and authentication: the MCP endpoint, `/oauth/*`, `/.well-known/*`, `llm.txt` /
`llm_full.txt`, the skills HTTP API, and custom `http.routes` (through the `http:ip-filter`
flow, before `auth: true` verification). Health and readiness probes (`/healthz`, `/readyz`,
`/health`) and `/metrics` (bearer-token protected) are exempt.

### Client IP Sources

| Runtime | Client IP |
| - | - |
| Node / Express | Socket peer; `::ffff:a.b.c.d` from a dual-stack socket matches IPv4 rules |
| Behind a declared proxy | `X-Forwarded-For` (see [Proxy Configuration](#proxy-configuration)) |
| Cloudflare Workers | `CF-Connecting-IP`, trusted only when running on Workers (Durable Objects included) |
| Deno | `info.remoteAddr` from `Deno.serve(handler)` |
| Bun | `server.requestIP(request)` from `Bun.serve({ fetch: handler })` |

### Supported IP Formats

| Format | Example |
| - | - |
| IPv4 address | `192.168.1.1` |
| IPv4 CIDR | `10.0.0.0/8` |
| IPv6 address | `2001:db8::1` |
| IPv6 CIDR | `2001:db8::/32` |
| IPv4-mapped IPv6 | `::ffff:192.168.1.1` (matched as `192.168.1.1`) |

### Proxy Configuration

When your server is behind a reverse proxy (Nginx, CloudFront, etc.), declare the proxy as
trusted so the client IP is read from `X-Forwarded-For` instead of the socket peer. This is
set through the environment:

```bash theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
FRONTMCP_TRUST_PROXY=true
FRONTMCP_TRUSTED_PROXY_DEPTH=2   # trust up to 2 proxy hops
```

The client is read that many hops back from the end of the chain, because callers can prepend
entries but only your own proxies append to it. A chain shorter than the configured depth is
ignored in favour of the socket peer.

<Note>
  `ipFilter.trustProxy` and `ipFilter.trustedProxyDepth` exist in the schema but are not read:
  client-IP extraction happens in the SDK context layer, before guard configuration is
  reachable, and setting either logs a startup warning. Use the environment variables above.
</Note>

<Warning>
  **`FRONTMCP_TRUST_PROXY` is only as good as your network boundary.** Trusting forwarded
  headers means trusting whoever can set them, so two things must hold:

  * **Every ingress path traverses the configured proxy chain.** If a caller can reach the
    origin directly -- a public origin IP, a peered VPC, a second ingress that skips the
    balancer -- they choose the whole `X-Forwarded-For` chain, and counting hops from its end
    just lands on an address they picked.
  * **The edge strips and rebuilds the forwarded headers.** The outermost proxy must discard
    any inbound `X-Forwarded-For` and `X-Real-IP` and write its own, so the only entries in the
    chain are ones your proxies appended.

    With depth `1` and no `X-Forwarded-For` at all, `X-Real-IP` is used -- the single-hop nginx
    convention. It is never consulted alongside a chain, because a caller can send both.
</Warning>

***

## App-Level Configuration

The `throttle` field in `@FrontMcp` configures all guard features at the app level:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@FrontMcp({
  name: 'production-server',
  throttle: {
    enabled: true,

    // Storage backend (defaults to in-memory)
    storage: { type: 'redis', redis: { config: { host: 'localhost', port: 6379 } } },
    keyPrefix: 'mcp:guard:',

    // Global limits (checked before per-entity)
    global: { maxRequests: 1000, windowMs: 60_000, partitionBy: 'ip' },
    globalConcurrency: { maxConcurrent: 50 },

    // Defaults for entities without explicit config
    defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: 'session' },
    defaultConcurrency: { maxConcurrent: 10 },
    defaultTimeout: { executeMs: 30_000 },

    // IP filtering
    ipFilter: {
      allowList: ['203.0.113.0/24'],
      denyList: ['192.0.2.1'],
      defaultAction: 'allow',
    },
  },
  tools: [SearchTool, ReportTool],
})
class ProductionApp {}
```

### Configuration Precedence

| Guard Type | Per-Entity Config | App Default | Fallback |
| - | - | - | - |
| Rate limit | `@Tool({ rateLimit })` | `throttle.defaultRateLimit` | No limit |
| Concurrency | `@Tool({ concurrency })` | `throttle.defaultConcurrency` | No limit |
| Timeout | `@Tool({ timeout })` | `throttle.defaultTimeout` | No timeout |
| IP filter | N/A (app-level only) | `throttle.ipFilter` | No filter |
| Global rate limit | N/A (app-level only) | `throttle.global` | No limit |

***

## Storage Backends

Guard supports multiple storage backends for distributed deployments.

### Memory (Development)

The default backend. Suitable for single-instance development. No configuration needed.

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
throttle: {
  enabled: true,
  // storage not set = in-memory
}
```

<Warning>
  In-memory storage does not persist across restarts and does not work with multiple server instances. Use Redis for production.
</Warning>

### Redis (Production)

For distributed rate limiting across multiple server instances:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
throttle: {
  enabled: true,
  storage: {
    type: 'redis',
    redis: {
      config: {
        host: 'redis.example.com',
        port: 6379,
        password: process.env.REDIS_PASSWORD,
        tls: true,
      },
    },
  },
}
```

`storage` is a `StorageConfig` from `@frontmcp/utils`, not the top-level `redis` shape: the backend is chosen by `type`, and its connection goes under the key of that name (`redis: { config }` or `redis: { url: process.env.REDIS_URL }`). A block without `type` (such as `{ provider: 'redis', host, port }`) is auto-detected from `REDIS_URL` / `REDIS_HOST` and otherwise runs in memory.

Redis enables pub/sub-based semaphore notifications for more efficient concurrency slot release detection.

### Vercel KV / Upstash

For serverless environments:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
throttle: {
  enabled: true,
  storage: {
    type: 'vercel-kv',
    vercelKv: {
      url: process.env.KV_REST_API_URL,
      token: process.env.KV_REST_API_TOKEN,
    },
  },
}
```

Upstash is `{ type: 'upstash', upstash: { url, token } }`.

### When the store is unreachable

Rate limits are a security control, so they **fail closed**. If the configured backend cannot be reached at startup, the server does not start: `createGuardManager` rejects with `GuardStorageUnavailableError` (`GUARD_STORAGE_UNAVAILABLE`), whose message names `throttle.storage`. This is the default in production, where `fallback` defaults to `'error'` (it defaults to `'memory'` otherwise).

The same applies if the backend goes away **while the server is running**: a rate-limited or concurrency-limited call is refused with the same `GuardStorageUnavailableError` (503, `GUARD_STORAGE_UNAVAILABLE`, a public error the client can read), not an internal error. The Redis client logs the connection error at a rate-limited interval instead of once per reconnect attempt.

To keep serving with per-instance counters while the store is down, opt in explicitly:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
throttle: {
  enabled: true,
  storage: {
    type: 'redis',
    redis: { url: process.env.REDIS_URL },
    fallback: 'memory', // each instance counts on its own while Redis is down, and goes back to Redis once it answers
  },
}
```

This differs from the top-level `redis` and `transport.persistence`, which fall back to in-memory storage with an error log.

### Key format

Keys are `<keyPrefix><entity>:<partition>:<kind>:…`, for example `mcp:guard:export_tickets:global:rl:1790722980000`. A trailing `:` on `keyPrefix` is dropped, because the storage namespace adds its own separator.

<Note>
  Before 1.8.6 the default prefix produced a double colon (`mcp:guard::export_tickets:…`). Counters written by an older instance are not read by a newer one, so during a rolling deploy to 1.8.6 each version counts separately until the old instances are gone, and a window's limit can briefly be exceeded. A custom `keyPrefix` without a trailing colon keeps its keys.
</Note>

***

## Error Handling

Guard throws specific error classes when limits are exceeded:

| Error Class | Code | HTTP Status | When Thrown |
| - | - | - | - |
| `RateLimitError` | `RATE_LIMIT_EXCEEDED` | 429 | Request exceeds rate limit |
| `ConcurrencyLimitError` | `CONCURRENCY_LIMIT` | 429 | No concurrency slot available |
| `QueueTimeoutError` | `QUEUE_TIMEOUT` | 429 | Queue wait time exceeded |
| `ExecutionTimeoutError` | `EXECUTION_TIMEOUT` | 408 | Execution exceeded deadline |
| `IpBlockedError` | `IP_BLOCKED` | 403 | Exported for your own code |
| `IpNotAllowedError` | `IP_NOT_ALLOWED` | 403 | Exported for your own code |
| `GuardStorageUnavailableError` | `GUARD_STORAGE_UNAVAILABLE` | 503 | `throttle.storage` is unreachable, at startup or while running |

A `tools/call` that hits a limit answers with an error result whose `_meta.code` is the code above, in production too.
The built-in `ipFilter` does not throw: it answers the HTTP request with 403 before any tool runs (see [Filter Precedence](#filter-precedence)).

***

## Agent Guard

Agents support the same guard options as tools:

```typescript theme={"theme":{"light":"snazzy-light","dark":"dark-plus"}}
@Agent({
  name: 'research-agent',
  description: 'Research assistant',
  rateLimit: { maxRequests: 10, windowMs: 60_000, partitionBy: 'userId' },
  concurrency: { maxConcurrent: 2 },
  timeout: { executeMs: 120_000 },
})
class ResearchAgent extends AgentContext {
  async execute(input: unknown) {
    // Agent execution with guard protection
  }
}
```

The agent flow follows the same stage ordering: `acquireQuota` → `acquireSemaphore` → `execute` (with timeout) → `releaseSemaphore` → `releaseQuota`.

***

## Configuration Reference

### `RateLimitConfig`

| Field | Type | Default | Description |
| - | - | - | - |
| `maxRequests` | `number` | *required* | Maximum requests allowed in the window |
| `windowMs` | `number` | `60000` | Time window in milliseconds |
| `partitionBy` | `PartitionKey` | `'global'` | Partition strategy for bucketing |

### `ConcurrencyConfig`

| Field | Type | Default | Description |
| - | - | - | - |
| `maxConcurrent` | `number` | *required* | Maximum simultaneous executions |
| `queueTimeoutMs` | `number` | `0` | Max wait time for a slot (0 = no wait) |
| `partitionBy` | `PartitionKey` | `'global'` | Partition strategy for bucketing |

### `TimeoutConfig`

| Field | Type | Default | Description |
| - | - | - | - |
| `executeMs` | `number` | *required* | Maximum execution time in milliseconds |

### `IpFilterConfig`

| Field | Type | Default | Description |
| - | - | - | - |
| `allowList` | `string[]` | `[]` | IPs or CIDR ranges to always allow |
| `denyList` | `string[]` | `[]` | IPs or CIDR ranges to always block |
| `defaultAction` | `'allow' \| 'deny'` | `'allow'` | Action when IP matches neither list, or no client IP is known |
| `trustProxy` | `boolean` | `false` | **Not read** (startup warning); set `FRONTMCP_TRUST_PROXY` |
| `trustedProxyDepth` | `number` | `1` | **Not read** (startup warning); set `FRONTMCP_TRUSTED_PROXY_DEPTH` |

### `GuardConfig` (App-Level)

| Field | Type | Default | Description |
| - | - | - | - |
| `enabled` | `boolean` | *required* | Enable or disable all guard features |
| `storage` | `StorageConfig` | in-memory | Storage backend (`{ type, redis \| vercelKv \| upstash, fallback? }`); fails closed unless `fallback: 'memory'` |
| `keyPrefix` | `string` | `'mcp:guard:'` | Prefix for all storage keys (a trailing `:` is dropped) |
| `global` | `RateLimitConfig` | — | Global rate limit for all requests |
| `globalConcurrency` | `ConcurrencyConfig` | — | Global concurrency limit |
| `defaultRateLimit` | `RateLimitConfig` | — | Default per-entity rate limit |
| `defaultConcurrency` | `ConcurrencyConfig` | — | Default per-entity concurrency |
| `defaultTimeout` | `TimeoutConfig` | — | Default per-entity timeout |
| `ipFilter` | `IpFilterConfig` | — | IP filtering configuration |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.