Guide
429 Too Many Requests: handle API rate limits with retries that work
A 429 is the politest error on the web: the server tells you what went wrong and, often, exactly how long to wait. Most clients ignore both halves of the message.
On this page
Key takeaways
- HTTP 429 Too Many Requests means you exceeded a rate limit or a quota; the request was refused, not processed.
- Wait for the Retry-After header when it is present; otherwise back off exponentially with random jitter, and cap the attempts.
- Tell a per-second 429 from a quota 429: the first clears within a second, the second will not clear by retrying.
- Every authenticated TickerLayer REST response carries X-RateLimit-Limit and X-RateLimit-Remaining, so a client can pace itself before it ever sees a 429.
- For live prices, one WebSocket connection replaces thousands of polling requests an hour.
429 Too Many Requests is the HTTP status code a server returns when a client has sent more requests than its rate limit allows in a given time window (RFC 6585). The request was refused, not processed, and nothing is wrong with it except its timing. The fix has two halves: retry correctly, by waiting for Retry-After or backing off exponentially with jitter, and then send fewer requests.
Market data clients hit 429s more than most, because the natural first design is to poll prices in a loop. The code below fixes the retries in Python and JavaScript. The lasting fix for live prices is to stream them from a stock WebSocket API and keep REST for history and lookups.
What an HTTP 429 response tells you
A useful 429 carries three pieces of information: which limit you hit, how long to wait, and what your budget is. TickerLayer sends all three. This is a per-second 429 for a key with a limit of 10 requests per second:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json; charset=utf-8
Retry-After: 1
X-RateLimit-Limit: 10
X-RateLimit-Remaining: 0
X-RateLimit-Policy: 10;w=1The response body
{
"error": "REST_RATE_LIMIT_EXCEEDED",1
"message": "rest rate limit exceeded",
"limitRps": 10,2
"retryAfterMs": 2503
}
errorWhich limit you hit. Branch on this code, never on the message text.limitRpsYour per-second budget, the same number asX-RateLimit-Limit.retryAfterMsMilliseconds until the current one-second window closes. More precise than the header.
Not every 429 is a pacing problem. The same status also reports an exhausted monthly quota, and the two call for opposite reactions:
| Per-second limit | Monthly quota | |
|---|---|---|
error | REST_RATE_LIMIT_EXCEEDED | REST_QUOTA_EXCEEDED |
| Cause | Too many requests inside one second | The account used its monthly allowance |
Retry-After header | Yes, in whole seconds | No |
| Does retrying help? | Yes, after the wait | No, not until the quota resets or the plan changes |
| Fix | Pace, back off, batch | Cut volume, stream instead of polling, or upgrade |
Two details from the rate limit reference change how you write the client. The per-second limit is checked first, so requests refused for pacing do not consume monthly quota. And the budget belongs to the account, so two services sharing one account share one budget. The full body shapes for every status are in the errors reference.
How servers count: four API rate limit algorithms
An API rate limit caps requests per unit of time, per key or per account. Servers count in one of four ways, and the one in use changes how bursts behave:
| Feature | How it counts | Bursts | What clients notice |
|---|---|---|---|
| Fixed window | Requests per calendar second or minute; the counter resets at the boundary | Up to twice the limit across a boundary | 429s cluster at the start of busy windows |
| Sliding window | Requests in the last N seconds, measured continuously | Smoothed out | Steady refusals under load, fewer edge effects |
| Token bucket | Tokens refill at a steady rate; each request spends one | Allowed up to the bucket size | Short bursts pass, sustained excess does not |
| Leaky bucket | Requests queue and drain at a fixed rate | Queued, not refused, until the queue fills | Latency rises before 429s appear |
The window length matters more than the algorithm. A limit of 600 requests per minute and a limit of 10 per second allow the same average, but the per-minute limit lets you burn the whole minute in the first two seconds and then refuses you for 58. Per-second limits punish bursts immediately and forgive them just as fast, which is why pacing works so well against them.
Retry correctly: exponential backoff with jitter
Exponential backoff doubles the wait after each failed attempt: 1 s, 2 s, 4 s, 8 s. Doubling alone has a flaw. If a hundred clients fail at the same instant, they also retry at the same instants, and each retry wave hits the server like the original burst. Random jitter spreads them out. The simplest version, full jitter, picks a random wait between zero and the exponential ceiling:
delay = random(0, min(cap, base × 2^attempt))with a server hint: delay = Retry-After + random(0, base)
- base
- The unit the ceiling doubles from, for example 0.5 s, so the first retry waits up to 1 s.
- attempt
- Which retry this is: 1, 2, 3 and so on.
- cap
- The longest you will wait between two tries, for example 20 s.
- Retry-After
- The server's own estimate. When it is present, it wins.
Retries reaching a recovering server
- Backoff without jitter
- Backoff with full jitter
requests per second
Four rules keep a retry loop from making things worse:
- Retry only what can succeed later: 429 and transient 5xx responses. A 400, 401, 403 or 404 fails the same way every time.
- Retry only idempotent requests. Every market data read is a
GET, so this is easy here; be careful with writes elsewhere. - Cap the attempts and the total time. Five attempts is plenty; after that, surface the error.
- Let the server hint win. If
Retry-Afteror a body field says when, use it, plus a little jitter.
Python: a retry wrapper you can paste
This wrapper uses requests. It reads the precise retryAfterMs from the body when present, falls back to Retry-After in either of its two formats (seconds or an HTTP date), and otherwise uses full jitter. A quota 429 raises at once, because waiting will not fix it.
import os
import random
import time
from email.utils import parsedate_to_datetime
import requests
BASE_URL = "https://api.tickerlayer.com"
RETRYABLE = {429, 500, 502, 503, 504}
session = requests.Session()
session.headers["x-api-key"] = os.environ["TICKERLAYER_API_KEY"]
class QuotaExceeded(Exception):
"""The monthly quota is used up. Retrying will not help until it resets."""
def body_of(resp):
try:
body = resp.json()
except ValueError:
return {}
return body if isinstance(body, dict) else {}
def server_wait(resp, body):
"""Seconds the server asked us to wait, or None if it did not say."""
if isinstance(body.get("retryAfterMs"), (int, float)):
return body["retryAfterMs"] / 1000 # more precise than the header
value = resp.headers.get("Retry-After")
if not value:
return None
if value.strip().isdigit():
return float(value)
try: # the HTTP-date form: "Wed, 21 Oct 2026 07:28:00 GMT"
return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
except (TypeError, ValueError):
return None
def get(path, params=None, max_attempts=5, base=0.5, cap=20.0):
"""GET a TickerLayer path, retrying 429 and transient 5xx responses."""
for attempt in range(1, max_attempts + 1):
resp = session.get(BASE_URL + path, params=params, timeout=10)
if resp.status_code not in RETRYABLE:
resp.raise_for_status() # 400, 401, 403, 404: fix the request instead
return resp.json()
body = body_of(resp)
if body.get("error") == "REST_QUOTA_EXCEEDED":
raise QuotaExceeded(f"monthly quota of {body.get('limit')} requests used up")
if attempt == max_attempts:
resp.raise_for_status()
hint = server_wait(resp, body)
if hint is not None:
delay = hint + random.uniform(0, base) # honour the server, de-synchronise a little
else:
delay = random.uniform(0, min(cap, base * 2 ** attempt)) # full jitter
print(f"{resp.status_code} on {path} (attempt {attempt}), retrying in {delay:.2f}s")
time.sleep(delay)
if __name__ == "__main__":
quote = get("/forex/quote/EURUSD")
print(quote["symbol"], quote["bid"], quote["ask"], quote["timestamp"])EURUSD 1.137233 1.137242 1790592920001To exercise the retry path without hammering a real API, we pointed the same function at a local test server that answered 429 (with retryAfterMs), 429 (with only the header) and 503 (with no hint) before succeeding. Each wait follows the rule for its case:
429 on /forex/quote/EURUSD (attempt 1), retrying in 0.27s
429 on /forex/quote/EURUSD (attempt 2), retrying in 1.32s
503 on /forex/quote/EURUSD (attempt 3), retrying in 1.67s
EURUSD 1.137233 1.137242 1790592920001JavaScript: the same logic with fetch
Node 18 and later ship fetch, so there is nothing to install. Save the file as retry.mjs so top-level await works.
const BASE_URL = "https://api.tickerlayer.com";
const API_KEY = process.env.TICKERLAYER_API_KEY;
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
class QuotaExceeded extends Error {}
// Milliseconds the server asked us to wait, or null if it did not say.
function serverWaitMs(res, body) {
if (typeof body.retryAfterMs === "number") return body.retryAfterMs;
const value = res.headers.get("retry-after");
if (!value) return null;
if (/^\d+$/.test(value.trim())) return Number(value) * 1000;
const at = Date.parse(value); // the HTTP-date form
return Number.isNaN(at) ? null : Math.max(0, at - Date.now());
}
async function get(path, { maxAttempts = 5, baseMs = 500, capMs = 20_000 } = {}) {
for (let attempt = 1; ; attempt++) {
const res = await fetch(BASE_URL + path, {
headers: { "x-api-key": API_KEY },
signal: AbortSignal.timeout(10_000),
});
const body = await res.json().catch(() => ({}));
if (res.ok) return body;
if (!RETRYABLE.has(res.status)) {
throw new Error(`${res.status} ${body.message ?? res.statusText}`); // fix the request
}
if (body.error === "REST_QUOTA_EXCEEDED") {
throw new QuotaExceeded(`monthly quota of ${body.limit} requests used up`);
}
if (attempt >= maxAttempts) throw new Error(`${res.status} after ${attempt} attempts`);
const hint = serverWaitMs(res, body);
const delay = hint !== null
? hint + Math.random() * baseMs // honour the server
: Math.random() * Math.min(capMs, baseMs * 2 ** attempt); // full jitter
console.warn(`${res.status} on ${path} (attempt ${attempt}), retrying in ${Math.round(delay)} ms`);
await sleep(delay);
}
}
const quote = await get("/forex/quote/EURUSD");
console.log(quote.symbol, quote.bid, quote.ask, quote.timestamp);EURUSD 1.13725 1.13726 1790592934000
429 on /forex/quote/EURUSD (attempt 1), retrying in 394 ms
429 on /forex/quote/EURUSD (attempt 2), retrying in 1222 ms
503 on /forex/quote/EURUSD (attempt 3), retrying in 1529 ms
EURUSD 1.137233 1.137242 1790592920001The Python stock prices tutorial builds a full client on the same pattern. Streams need their own version of backoff for reconnects, covered in WebSocket reconnect.
Stop hitting the limit: budget and pace your requests
Retries treat the symptom. Most 429s in market data come from a handful of request patterns, and each has a cheaper shape that returns the same data:
| Pattern | Requests | Cheaper shape | Requests |
|---|---|---|---|
| Poll 20 symbols every 2 s | 36,000 an hour | Stream them over one WebSocket | 1 subscribe message |
| Quote, last trade and previous close per symbol | 3 per symbol | GET /{asset}/snapshot/{symbol} | 1 per symbol |
| Open or closed, for 25 markets | 25 | GET /markets/status/bulk?items= | 1 |
| Fundamentals for 100 US stocks | 100 | GET /fundamentals/stocks?symbols= (up to 100) | 1 |
| A year of 1-minute bars for one US stock | Hundreds of pages at the default 500 | limit=5000 per page | A tenth as many |
| Polling through the weekend | All of them wasted | Read next_open from /markets/status, then sleep | 1 |
The market status routes, including the bulk form and next_open, are documented in the market status reference. For polling you cannot avoid, work out the budget before you write the loop:
Inside the per-second limit, pace rather than burst. X-RateLimit-Limit tells you the budget, so a client can space its requests evenly and never see a 429 at all. This pacer keeps 20% headroom and is safe to share across threads:
import os
import threading
import time
import requests
BASE_URL = "https://api.tickerlayer.com"
session = requests.Session()
session.headers["x-api-key"] = os.environ["TICKERLAYER_API_KEY"]
class Pacer:
"""Spaces requests evenly so a burst never exceeds the per-second budget."""
def __init__(self, per_second, headroom=0.8):
self.interval = 1.0 / (per_second * headroom)
self.next_at = time.monotonic()
self.lock = threading.Lock()
def wait(self):
with self.lock:
now = time.monotonic()
start = max(now, self.next_at)
self.next_at = start + self.interval
time.sleep(max(0.0, start - now))
first = session.get(f"{BASE_URL}/forex/quote/EURUSD", timeout=10)
first.raise_for_status()
limit = int(first.headers.get("X-RateLimit-Limit", "10"))
pacer = Pacer(limit)
print(f"limit {limit}/s, one request every {pacer.interval * 1000:.0f} ms")
for symbol in ["GBPUSD", "USDJPY", "AUDUSD"]:
pacer.wait()
resp = session.get(f"{BASE_URL}/forex/quote/{symbol}", timeout=10)
resp.raise_for_status()
quote = resp.json()
print(symbol, quote["bid"], quote["ask"], "remaining:", resp.headers["X-RateLimit-Remaining"])limit 10/s, one request every 125 ms
GBPUSD 1.325809 1.325879 remaining: 8
USDJPY 157.0403 157.0573 remaining: 7
AUDUSD 0.701173 0.701233 remaining: 6Cache what does not change between polls. Symbol lists, market sessions and holiday calendars change rarely, so fetch them once a day. For prices, keep one server-side cache per symbol for as long as your screen can tolerate, and serve every viewer from it instead of calling the API once per page view. When the whole design is a live screen, stop polling: WebSocket vs REST works through when the switch pays off.
Debug a 429 in production
- Read the bodyBranch on
error.REST_QUOTA_EXCEEDEDis a budget problem;REST_RATE_LIMIT_EXCEEDEDis a pacing problem. - Log the headersRecord
X-RateLimit-Remainingon every response. A value that hits 0 at the same second of every minute points at a scheduled job. - Find the burstLook for cron jobs firing on the minute, retries without jitter, and several services sharing one account.
- Check the monthThe dashboard shows usage against the monthly quota; the response headers do not.
- Change the shapeBatch, cache, pace or stream, then delete any retry loop you added as a workaround.
Before you ship a client
- Retries on 429 and transient 5xx only, with a cap on attempts.
Retry-Afterhonoured in both the seconds and the HTTP-date form.- Full jitter on every computed backoff.
- Quota 429s raised as an alert, not retried.
- Requests paced from
X-RateLimit-Limit, with headroom. - Live prices streamed rather than polled once you watch more than a few symbols.
- Polling paused while the market is closed.
Rate limits are one part of a healthy integration; symbols, units and timestamps are the others, covered in market data API explained.
Questions
What does 429 Too Many Requests mean?
The server refused your request because you exceeded a rate limit or a quota. The request itself is fine; send it again after the wait the server indicates.
How long should I wait after a 429 error?
As long as the Retry-After header says, either in seconds or until the HTTP date it gives. With no header, back off exponentially with random jitter, starting around half a second.
Is HTTP 429 a client error or a server error?
A client error: it is in the 4xx range because the client sent too many requests. Unlike most 4xx codes, it is safe to retry after waiting.
Do rate-limited requests count against my quota?
On TickerLayer, no. The per-second limit is checked first, so a request refused with REST_RATE_LIMIT_EXCEEDED does not consume monthly quota.
What is the difference between 429 and 503?
A 429 says you are sending too much; a 503 says the server cannot serve requests right now. Both can carry Retry-After, and both deserve backoff with jitter.
How do I fix 429 errors in Python?
Wrap requests in a retry loop that honours Retry-After, backs off with full jitter, caps attempts and does not retry quota errors. Then reduce volume by batching, caching or streaming.