Skip to main content
Limits are per key, not per IP or per account. Three independent token buckets refill continuously: A bucket that starts full can emit burst + sustained requests inside one second, so a read client can briefly reach 300 before throttling. Write burst equals the sustained rate — no extra spike headroom on places. Cost is 1 token per HTTP request. A 20-order batch costs 1 token of the matching bucket, not 20. Batches are capped at 20 items (400 batch_too_large before any token is charged). The cancel bucket is always 2 × the effective write rate. It is not independently overridable. Cancels do not share the write bucket, so a round-end ladder pull cannot 429 behind places.

Volume-scaled write rate

Write rate is linear in your account’s trailing 30 UTC calendar days of fill volume (SUM(qty) of fills, maker and taker, $1/share). A short history is scored on the days that exist — it is not annualized:
The floor is 20/s so a brand-new account can still trade. Volume buys market-making headroom. The cap is 400/s so rate + burst stays under the per-user open-order kill switch. There are no named usage tiers. Prefer GET /v1/account/limits over hard-coding the table — that response is the grant actually attached to your key.

When you exceed a bucket

You get HTTP 429 with rate_limited, and the response body carries retry_after_ms in details:
Back off for that long and retry. A 429 is a soft throttle with no penalty — you are not penalised for hitting it, and the next window is clean.
retry_after_ms is in the response body, not in a Retry-After header. HTTP client libraries that back off automatically on Retry-After will not see it. Read it from the body yourself.
WebSocket subscribe commands that exhaust the read bucket return wscode 10 instead of HTTP 429. Back off the same way.

A 429 is not a suspension

Worth separating two different kinds of “no”:
  • 429 rate_limited, a throttle. Back off, carry on.
  • Suspension, a risk-management action against your account. New orders are refused until it is lifted. Cancels still go through. This is not something you retry your way out of.
If writes start failing with something other than 429, stop and read the error code rather than retrying harder.

Staying under

  • Poll market data on an interval rather than in a tight loop. Rounds last 60 seconds; polling faster than a few times a second gains you nothing. Use the WebSocket for anything latency-sensitive.
  • Batch order operations where you can — one HTTP request is one token.
  • Use one key per process. Two processes sharing a key share its buckets.

Discover your key’s limits

Rates, bursts, and knobs are JSON numbers. trailing_volume_30d is a decimal string. The response has no as_of_seq — it is answered from the key record, not a stream fold.

Higher limits

Write rate scales automatically with trailing 30-day fill volume — see the formula above. Manual per-key grants exist and are never overwritten by the volume formula. Do not shard across keys to dodge the limit: two keys mean two buckets, but the open-order kill switch is per user.