Guides
Rate limits
Limits are counted per API key and per model, in fixed UTC windows. Going over returns 429 with a Retry-After header.
What is limited
| Limit | Counts | Window |
|---|---|---|
| RPM | Requests | Fixed 1-minute window aligned on the UTC minute (12:03:00 to 12:03:59) |
| TPM | Tokens (estimated when the request arrives, corrected after it completes) | Same 1-minute window |
| RPD | Requests | UTC calendar day, reset at 00:00 UTC |
- Counters are kept per API key and per model: 30 requests to
openai/gpt-4oand 30 toopenai/gpt-4o-miniin the same minute use two separate budgets. - TPM applies to token-priced models. Video models have RPM and RPD only.
- A request counts as soon as it passes authentication, even if it fails afterwards (for example with
402). rodiumai/smartmakes a router call first, which counts against the router model (openai/gpt-4o-miniby default) as well as the selected model.
Which values apply
- The model's own limits, published in
GET /v1/modelsunderrodiumai_rate_limits(rpm,tpm,rpd). For exampleopenai/gpt-4olists 30 RPM, 1,000,000 TPM and 2,000 RPD at the time of writing. - Where the model sets no RPM or TPM: the key's own RPM/TPM, when RodiumAI has configured one for that key.
- Otherwise the defaults: 60 requests and 60,000 tokens per minute. RPD only applies when the model defines it.
- For team (organization) accounts, RodiumAI can then scale these values with a multiplier or set fixed values, globally or per model.
A null field in rodiumai_rate_limits means the model sets no value for that dimension and the next rule applies.
How tokens are counted
When the request arrives, its tokens are estimated (prompt plus max_tokens) and added to the minute's TPM counter, at most 16,000 tokens per request. After the call, if the real usage was higher than that estimate, the difference (within the same 16,000 cap) is added to the current minute. Setting a realistic max_tokens therefore also keeps you further from the TPM limit.
When you hit a limit
…Retry-Afteris in seconds:60for RPM and TPM, and the time left until 00:00 UTC (at least 60) for RPD.- There are no
X-RateLimit-*headers: track your own request rate, or read the limits fromGET /v1/models. - Nothing is billed for a rate-limited request.
- A
429can also come from the model provider; it has the same envelope and aRetry-Afterheader, and is handled the same way.
Staying under the limits
- Wait for
Retry-After, then retry with jitter; cap the number of attempts. A ready-made loop is in Errors & retries. - Spread bursts: a queue with a fixed concurrency is easier to keep under RPM than parallel bursts.
- Split traffic across models where quality allows; every model has its own budget.
- Need higher limits for production? Contact support with your expected requests and tokens per minute.