Guides
Rate limits
How fast a key may call the API, the headers that say how much is left, and what to do at a limit.
Every response says which limit it was counted against and how much of it is left.
The limits
| Limit | Default |
|---|---|
| Reads, per key | 600 a minute |
| Writes, per key | 120 a minute |
| Chat messages, per key | 1,200 a minute |
| Chat messages, per conversation | 20 a minute |
| Each of the three above, for all of an organization's keys together | 4 × the per-key number |
| Calls at the same time, per organization | 10 |
| Calls a day, per organization | 2,000 |
| Recording links, per key | 60 a minute and 2,000 a day |
| Test calls, per test key | 10 a minute |
The deployment's operator can change an organization's limits.
The headers
RateLimit-Policy: "writes";q=120;w=60
RateLimit: "writes";r=87;t=34
q is the limit, w the window in seconds, r what is left in this window and t the seconds until
it starts again. The format follows the IETF's RateLimit header draft. A 429 also has
Retry-After, in seconds.
At a limit
| Code | What to do |
|---|---|
rate_limited |
Wait for Retry-After, then send the same request again. |
concurrency_limit_reached |
Wait for calls to finish; GET /v1/me shows how many are in progress. Do not keep retrying. |
daily_limit_reached |
The day's calls are used. Try again after midnight UTC. |
limits_unavailable |
Something that costs money could not be counted safely just now. Wait for Retry-After. |
A retry is never refused for the request it repeats: a request sent again with its
Idempotency-Key gets its first answer back before any limit is counted.