Rate Limiting Specification

Overview

The Passport for AI Agents rate limits on a sliding window to bound abuse. The
counters are KV-backed and approximate, not quotas: see "When the counter cannot
be written" below before treating any number here as a hard ceiling.

Rate Limits

Four tiers, defined in RATE_LIMIT_TIERS in
functions/utils/rate-limiting/rate-limit.ts. The numbers below are the tier
defaults, which apply when no environment variable and no per-route value is set.

Tier Default per minute Declared environment value Endpoints
verify 10,000 VERIFY_RPM: 60 dev, 500 preview, 1000 production /api/verify/*, passport reads and writes
admin 480 ADMIN_RPM: 480 everywhere /api/admin/*, /api/metrics
org 300 ORG_RPM: 300 everywhere /api/orgs/, /api/org/
policy_verify 100,000 POLICY_RATE_LIMIT_PER_MINUTE: 100,000 preview, 200,000 production /api/verify/policy/*

These environment values come from wrangler.toml; they do not establish what
a running deployment receives. A non-browser passport verification request to
api.aport.io on September 30, 2026 returned HTTP 200 and
X-RateLimit-Limit: 10000, despite the named production section declaring 1000.
The deployed policy-verify override was not independently measured. These
allowances are not evidence of sustained throughput; that requires a load test.

Authentication protection is a separate limit

AuthRateLimiter in functions/utils/security/security-enhancements.ts is a
separate in-memory admission and failure counter. It has a threshold of 50, a
15-minute idle window, and a 30-minute block. Successful JWT authentication
resets it. Successful API-key authentication now performs the same reset;
previously, the 51st valid API-key verification could lock the caller out.
Increasing VERIFY_RPM or POLICY_RATE_LIMIT_PER_MINUTE could not fix that.

Requests are counted before credential lookup. Each admitted request holds a
reservation until authentication returns or throws. Success clears completed
failures while retaining the charge for every other pending lookup; neither
success nor idle-window/cooldown expiry can erase pending work. The first
request after cooldown expiry is also counted (the old reset missed it).
Some JWT rejection paths also record a failure, so they reach the threshold
after 25 rejected requests; paths counted only at admission reach it after 50.
API-key scope denials and service errors retain their existing counting.
The existing development exception raises the threshold to 500 and reduces
the block to five minutes for requests classified as local.

This counter is a per-worker map keyed by client IP and Accept-Language. It is
best-effort protection, not a distributed quota or a per-passport allowance.
Callers sharing that fingerprint still share state. More than 50 concurrent
lookups hit its admission cap even if their credentials will prove valid.
Saturation returns a transient 429 with Retry-After: 1; it does not start a
30-minute failed-credential cooldown. Completed failures still trigger the
normal cooldown. Per-credential quota isolation is separate work.

Policy verification, BaseApiHandler, and manual-auth routes share a response
formatter that preserves authentication lockouts as
HTTP 429, with error: "auth_rate_limit_exceeded", retry_after in seconds,
and a matching Retry-After header. This is a service refusal before policy
evaluation, so clients must not interpret it as an allow or deny decision.
A pre-tool hook that fails closed on evaluation errors will still stop the
tool call, including when policy denials are configured as warnings.

Which number a route gets

Most specific first:

  1. The number the route declares as rateLimitRpm in its ApiHandlerConfig.

47 routes declare one, from 40 on orgs/[id]/orgkey.ts to 480 on
verify/decisions/get/[decision_id].ts.

  1. The tier's environment variable.
  2. The tier default from the table above.

So the deployed limit for a named route is usually its own declared number, not
the tier value. GET /api/orgs/:id/members is 240 a minute even though ORG_RPM
is 30.

What this deploy changes about limits

Two things happen at once, and they pull in opposite directions.

Declared per-route numbers become effective. The environment variable used to
win over the number written at the endpoint, so 39 of 47 declared values were
dead. They now take effect. Separately, five routes were passing their config to a
constructor parameter that BaseApiHandler discards and then calling execute()
with no argument, so they ran on the verify tier at VERIFY_RPM rather than the
tier and number they declared.

Every route limit was then widened about 4x, and the tier defaults with it:
org 30 to 300, admin 100 to 480, and each declared rateLimitRpm multiplied
by four and capped at 480. This is a deliberate early-stage trade: a 429 shown to
a real customer costs more right now than the abuse a wider window admits.
verify (10,000) and policy_verify (100,000) are unchanged; they were already
far above anything the widening reaches.

Nothing is set to 500 or more, and that boundary is not stylistic.
UNCOUNTABLE_ALLOW_FLOOR_PER_MINUTE is 500 and the comparison is inclusive, so a
route configured at 500 fails open when its KV counter write saturates, which
is exactly the condition the limit exists to handle. 480 buys the headroom while
keeping every one of these routes failing closed. A route needing more than 480
should be argued for by name rather than nudged over the line.

Net effect, measured against what was actually enforced before this branch:

Routes Change
Loosened or unchanged 35 org routes gain the most, from an effective 30 to 120-480
Still tighter 12 verify-tier routes at ~1,000 before, now 480

The 12 that end up tighter are all on the verify tier, where the old effective
number was VERIFY_RPM (1,000 in production) because their declared values were
being discarded: passports/create.ts, passports/list.ts,
passports/[agent_id]/index.ts, passports/[agent_id]/webhooks.ts (both
methods), passports/[template_id]/instances-list.ts,
passports/by-slug/[slug]/index.ts, verify/decisions/get/[decision_id].ts,
decisions.ts, and the three claim/* routes. Each drops from about 1,000 to
480, roughly halving rather than the 8x cut they would have taken at their
pre-widening declared values.

The four webhook endpoints also begin requiring authentication, and
passports/:template_id/webhooks begins checking that the caller owns the
passport, which it never did.

What a counter measures

For routes that go through BaseApiHandler, one counter per tier, per route, per
client. The key is rate_limit::::, where
the route pattern has its dynamic segments masked (/api/orgs/:id/orgkey, not
/api/orgs/org_abc/orgkey).

Both halves matter. Without the route in the key, every route in a tier counts
against one number, so calls to a 240-a-minute route exhaust a 10-a-minute route
on the same tier. With the raw path instead of the masked pattern, each id gets
its own counter and a caller cycling ids has no limit at all.

This shape comes from BaseApiHandler.validateRateLimit, and not every route
goes through it. A handful build a limiter directly from the factories and call
checkLimit(clientIP), which produces rate_limit:: with no
route segment. Those routes share one counter per tier per IP:
admin/audit/[agent_id].ts, admin/issue-org-key.ts, admin/status.ts and
admin/update.ts all share the admin one, so calls to any of them consume the
same allowance. /api/verify/policy/* likewise keys on IP plus a request
fingerprint rather than the route.

That is the pre-existing shape for those call sites, not a regression, and moving
them onto the handler path is a separate change: they would each pick up the
handler's method and auth validation along with its counter key.

Browser passport reads

/api/verify/:agent_id has a browser fast path: common browser user agents and
conditional cached requests skip the KV rate-limit check. The endpoint is used
for public passport reads from web pages, so the verify tier is enforced for
non-browser clients such as SDKs, CLIs, bots and server-to-server traffic, not
for every browser-shaped request. A browser response can still include
rate-limit headers; that does not prove the request consumed the allowance.

Headers

Rate Limit Headers

Counted rate-limited responses return the following headers:

X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 999
X-RateLimit-Reset: 1640995200
X-RateLimit-Window: 60

Header Descriptions

  • X-RateLimit-Limit: Maximum requests allowed per window
  • X-RateLimit-Remaining: Requests remaining in current window
  • X-RateLimit-Reset: Unix timestamp when the window resets
  • X-RateLimit-Window: Window size in seconds

Rate Limit Exceeded Response

When rate limit is exceeded, the API returns:

{
  "error": "Rate limit exceeded",
  "message": "Too many requests. Please try again later.",
  "retryAfter": 30,
  "limit": 60,
  "remaining": 0,
  "resetTime": "2024-01-01T00:00:00Z"
}

Status Code: 429 Too Many Requests

Implementation Details

Sliding Window Algorithm

  • Uses Cloudflare KV for distributed rate limiting
  • Window slides every second
  • Counters are automatically cleaned up after window expires

Client Identification

  • Primary: Client IP address
  • Fallback: User-Agent + IP combination for edge cases

Configuration

Rate limits are configurable via environment variables. Each falls back to its
tier default when unset, not to 60:

  • VERIFY_RPM: verify tier (tier default 10,000)
  • ADMIN_RPM: admin tier (tier default 100)
  • ORG_RPM: org tier (tier default 30)
  • POLICY_RATE_LIMIT_PER_MINUTE: policy verification tier (tier default 100,000)

A route that declares its own rateLimitRpm ignores these; see "Which number a
route gets" above.

When the counter cannot be written

Counters live in KV, which sustains roughly one write per second per key, so a
single key counts about 60 requests a minute. A limit above that cannot be
enforced by this path, and a failed counter write is handled by the size of the
limit in force:

  • At or above 500 a minute, and on a tier whose whenUncountable is

"allow" (verify and policy_verify), the request is admitted and
NOT ENFORCED is logged. Capping an enterprise tier at 60 would refuse
legitimate traffic. 500 is the floor because it is the lowest
environment-configured tier ceiling in wrangler.toml (preview VERIFY_RPM);
the largest per-route rateLimitRpm is 480, so the two populations do not
overlap.

  • Otherwise the request is denied. That covers every limit below 500 and every

limit on the admin and org tiers whatever its size. The limit exists to
bound abuse, and a write to a counter key fails precisely when that key is
being hammered.

A read failure is handled the other way: if the count cannot be read at all,
the request is admitted with a full allowance. A KV disruption lifts the limit
rather than refusing traffic.

No limit here is exact, at any size. The limiter reads the count and writes it
back, so concurrent requests can read the same number and write the same
increment. The 30-a-minute organisation tier undercounts a burst for the same
reason the 200,000 policy tier does; what differs between the tiers is only the
write-failure behaviour above. Exact enforcement needs an atomic counter.
Cloudflare's rate limiting binding is declared under [[unsafe.bindings]], which
Pages rejects, so every deployment here takes the KV path.

Best Practices

For API Consumers

  1. Respect Rate Limits: Check X-RateLimit-Remaining header
  2. Implement Backoff: Use exponential backoff when rate limited
  3. Cache Responses: Use ETag headers for efficient caching
  4. Monitor Usage: Track your API usage patterns

For Developers

  1. Test Rate Limits: Include rate limiting in integration tests
  2. Handle 429 Responses: Implement proper retry logic
  3. Monitor Headers: Log rate limit headers for debugging
  4. Optimize Requests: Batch requests when possible

Examples

Successful Request

curl -i "https://api.aport.io/api/verify/ap_a2d10232c6534523812423eec8a1425c"

HTTP/1.1 200 OK
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 999
X-RateLimit-Reset: 1640995200
X-RateLimit-Window: 60
Content-Type: application/json

{
  "agent_id": "ap_a2d10232c6534523812423eec8a1425c",
  "status": "active",
  "owner": "AI Research Lab"
}

Rate Limited Request

curl -i "https://api.aport.io/api/verify/ap_a2d10232c6534523812423eec8a1425c"

HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1640995200
X-RateLimit-Window: 60
Retry-After: 30
Content-Type: application/json

{
  "error": "Rate limit exceeded",
  "message": "Too many requests. Please try again later.",
  "retryAfter": 30
}

Monitoring

Rate limiting metrics are available via the /api/metrics endpoint:

{
  "rateLimiting": {
    "verify": {
      "totalRequests": 1250,
      "rateLimitedRequests": 15,
      "rateLimitPercentage": 1.2
    },
    "admin": {
      "totalRequests": 340,
      "rateLimitedRequests": 2,
      "rateLimitPercentage": 0.6
    }
  }
}