# Rate limits

> How fast a key may call the API, the headers that say how much is left, and what to do at a limit.

Every response says which limit it was counted against and how much of it is left.

## The limits

| Limit | Default |
|---|---|
| Reads, per key | 600 a minute |
| Writes, per key | 120 a minute |
| Chat messages, per key | 1,200 a minute |
| Chat messages, per conversation | 20 a minute |
| Each of the three above, for all of an organization's keys together | 4 × the per-key number |
| Calls at the same time, per organization | 10 |
| Calls a day, per organization | 2,000 |
| Recording links, per key | 60 a minute and 2,000 a day |
| Test calls, per test key | 10 a minute |

The deployment's operator can change an organization's limits.

## The headers

```text
RateLimit-Policy: "writes";q=120;w=60
RateLimit: "writes";r=87;t=34
```

`q` is the limit, `w` the window in seconds, `r` what is left in this window and `t` the seconds until
it starts again. The format follows the IETF's RateLimit header draft. A `429` also has
`Retry-After`, in seconds.

## At a limit

| Code | What to do |
|---|---|
| `rate_limited` | Wait for `Retry-After`, then send the same request again. |
| `concurrency_limit_reached` | Wait for calls to finish; `GET /v1/me` shows how many are in progress. Do not keep retrying. |
| `daily_limit_reached` | The day's calls are used. Try again after midnight UTC. |
| `limits_unavailable` | Something that costs money could not be counted safely just now. Wait for `Retry-After`. |

**A retry is never refused for the request it repeats**: a request sent again with its
`Idempotency-Key` gets its first answer back before any limit is counted.
