All pages

Guides

Rate limits

How fast a key may call the API, the headers that say how much is left, and what to do at a limit.

View as Markdown

Every response says which limit it was counted against and how much of it is left.

The limits

Limit Default
Reads, per key 600 a minute
Writes, per key 120 a minute
Chat messages, per key 1,200 a minute
Chat messages, per conversation 20 a minute
Each of the three above, for all of an organization's keys together 4 × the per-key number
Calls at the same time, per organization 10
Calls a day, per organization 2,000
Recording links, per key 60 a minute and 2,000 a day
Test calls, per test key 10 a minute

The deployment's operator can change an organization's limits.

The headers

RateLimit-Policy: "writes";q=120;w=60
RateLimit: "writes";r=87;t=34

q is the limit, w the window in seconds, r what is left in this window and t the seconds until it starts again. The format follows the IETF's RateLimit header draft. A 429 also has Retry-After, in seconds.

At a limit

Code What to do
rate_limited Wait for Retry-After, then send the same request again.
concurrency_limit_reached Wait for calls to finish; GET /v1/me shows how many are in progress. Do not keep retrying.
daily_limit_reached The day's calls are used. Try again after midnight UTC.
limits_unavailable Something that costs money could not be counted safely just now. Wait for Retry-After.

A retry is never refused for the request it repeats: a request sent again with its Idempotency-Key gets its first answer back before any limit is counted.