What the status codes mean
HTTP 429 (Too Many Requests) means you've exceeded a rate limit. The server is healthy but asking you to slow down. HTTP 503 (Service Unavailable) means the upstream provider is temporarily overloaded or down — the issue is on their side, not yours.
Both are recoverable with the right strategy. Neither means your code is wrong.
Why AI APIs hit rate limits
Providers enforce rate limits to protect their infrastructure. The limits are usually based on one or more of: requests per minute (RPM), tokens per minute (TPM), requests per day (RPD), or concurrent connections.
Free-tier models tend to have the tightest limits. Paid models have higher caps but still enforce them. A gateway like AnyRouter adds another layer of rate limiting on top of the upstream provider's limits.
Practical fixes
- Add exponential backoff on 429s. Wait 1s before the first retry, 2s before the second, 4s before the third, and so on up to a reasonable max (60s).
- Respect the Retry-After header if the API provides one — it tells you exactly how long to wait.
- Reduce request frequency. Batch requests, switch to a faster model for non-critical tasks, or spread load across multiple API keys if allowed.
- Check whether the error is a 429 (your fault) or a 503 (upstream's fault). 503s usually resolve within minutes; 429s require you to change your request rate.
- Monitor usage in your dashboard so you can see when you're approaching limits before you hit them.
How AnyRouter handles this
AnyRouter automatically fails over to an alternative upstream when one returns a 429, 401, or 5xx. The retry happens server-side, so your client only sees the final successful response — no special error handling needed on your side.
If you're using the shared-key pool, rate limits are distributed across multiple provider keys, which raises the effective limit compared to a single key.
Monitoring and observability
Track your request patterns, error rates, and model-level latency at chmonitor.dev. Seeing which models are returning errors helps you decide whether to switch models, increase limits, or adjust your request pattern.
Route through AnyRouter with automatic failover and monitor everything at chmonitor.dev.
Sign up free