Guide

429 Too Many Requests: handle API rate limits with retries that work

A 429 is the politest error on the web: the server tells you what went wrong and, often, exactly how long to wait. Most clients ignore both halves of the message.

On this page
  1. What an HTTP 429 response tells you
  2. How servers count: four API rate limit algorithms
  3. Retry correctly: exponential backoff with jitter
  4. Python: a retry wrapper you can paste
  5. JavaScript: the same logic with fetch
  6. Stop hitting the limit: budget and pace your requests
  7. Debug a 429 in production
  8. Questions

Key takeaways

  • HTTP 429 Too Many Requests means you exceeded a rate limit or a quota; the request was refused, not processed.
  • Wait for the Retry-After header when it is present; otherwise back off exponentially with random jitter, and cap the attempts.
  • Tell a per-second 429 from a quota 429: the first clears within a second, the second will not clear by retrying.
  • Every authenticated TickerLayer REST response carries X-RateLimit-Limit and X-RateLimit-Remaining, so a client can pace itself before it ever sees a 429.
  • For live prices, one WebSocket connection replaces thousands of polling requests an hour.

429 Too Many Requests is the HTTP status code a server returns when a client has sent more requests than its rate limit allows in a given time window (RFC 6585). The request was refused, not processed, and nothing is wrong with it except its timing. The fix has two halves: retry correctly, by waiting for Retry-After or backing off exponentially with jitter, and then send fewer requests.

Market data clients hit 429s more than most, because the natural first design is to poll prices in a loop. The code below fixes the retries in Python and JavaScript. The lasting fix for live prices is to stream them from a stock WebSocket API and keep REST for history and lookups.

What an HTTP 429 response tells you

A useful 429 carries three pieces of information: which limit you hit, how long to wait, and what your budget is. TickerLayer sends all three. This is a per-second 429 for a key with a limit of 10 requests per second:

Response headers (documented shape)HTTP
HTTP/1.1 429 Too Many Requests
Content-Type: application/json; charset=utf-8
Retry-After: 1
X-RateLimit-Limit: 10
X-RateLimit-Remaining: 0
X-RateLimit-Policy: 10;w=1

The response body

{
  "error": "REST_RATE_LIMIT_EXCEEDED",1
  "message": "rest rate limit exceeded",
  "limitRps": 10,2
  "retryAfterMs": 2503
}
  1. errorWhich limit you hit. Branch on this code, never on the message text.
  2. limitRpsYour per-second budget, the same number as X-RateLimit-Limit.
  3. retryAfterMsMilliseconds until the current one-second window closes. More precise than the header.
The per-second 429 as documented in the errors reference. Retry-After rounds the wait up to whole seconds; the body gives the exact milliseconds.

Not every 429 is a pacing problem. The same status also reports an exhausted monthly quota, and the two call for opposite reactions:

Per-second limitMonthly quota
errorREST_RATE_LIMIT_EXCEEDEDREST_QUOTA_EXCEEDED
CauseToo many requests inside one secondThe account used its monthly allowance
Retry-After headerYes, in whole secondsNo
Does retrying help?Yes, after the waitNo, not until the quota resets or the plan changes
FixPace, back off, batchCut volume, stream instead of polling, or upgrade
Both arrive as HTTP 429. Read the body before you decide to retry.

Two details from the rate limit reference change how you write the client. The per-second limit is checked first, so requests refused for pacing do not consume monthly quota. And the budget belongs to the account, so two services sharing one account share one budget. The full body shapes for every status are in the errors reference.

How servers count: four API rate limit algorithms

An API rate limit caps requests per unit of time, per key or per account. Servers count in one of four ways, and the one in use changes how bursts behave:

FeatureHow it countsBurstsWhat clients notice
Fixed windowRequests per calendar second or minute; the counter resets at the boundaryUp to twice the limit across a boundary429s cluster at the start of busy windows
Sliding windowRequests in the last N seconds, measured continuouslySmoothed outSteady refusals under load, fewer edge effects
Token bucketTokens refill at a steady rate; each request spends oneAllowed up to the bucket sizeShort bursts pass, sustained excess does not
Leaky bucketRequests queue and drain at a fixed rateQueued, not refused, until the queue fillsLatency rises before 429s appear
A one-second window, which is what TickerLayer uses, recovers quickly: a pacing refusal never asks you to wait more than one second.

The window length matters more than the algorithm. A limit of 600 requests per minute and a limit of 10 per second allow the same average, but the per-minute limit lets you burn the whole minute in the first two seconds and then refuses you for 58. Per-second limits punish bursts immediately and forgive them just as fast, which is why pacing works so well against them.

Retry correctly: exponential backoff with jitter

Exponential backoff doubles the wait after each failed attempt: 1 s, 2 s, 4 s, 8 s. Doubling alone has a flaw. If a hundred clients fail at the same instant, they also retry at the same instants, and each retry wave hits the server like the original burst. Random jitter spreads them out. The simplest version, full jitter, picks a random wait between zero and the exponential ceiling:

delay = random(0, min(cap, base × 2^attempt))with a server hint: delay = Retry-After + random(0, base)

base
The unit the ceiling doubles from, for example 0.5 s, so the first retry waits up to 1 s.
attempt
Which retry this is: 1, 2, 3 and so on.
cap
The longest you will wait between two tries, for example 20 s.
Retry-After
The server's own estimate. When it is present, it wins.
Attempt 3 with a 0.5 s base and a 20 s cap: a random wait between 0 and 4 s.

Retries reaching a recovering server

  • Backoff without jitter
  • Backoff with full jitter

requests per second

Illustrative, x axis in seconds: 100 clients and a server that can take about 30 requests a second. Synchronized retries arrive as spikes above capacity; jittered retries arrive under it.

Four rules keep a retry loop from making things worse:

  1. Retry only what can succeed later: 429 and transient 5xx responses. A 400, 401, 403 or 404 fails the same way every time.
  2. Retry only idempotent requests. Every market data read is a GET, so this is easy here; be careful with writes elsewhere.
  3. Cap the attempts and the total time. Five attempts is plenty; after that, surface the error.
  4. Let the server hint win. If Retry-After or a body field says when, use it, plus a little jitter.

Python: a retry wrapper you can paste

This wrapper uses requests. It reads the precise retryAfterMs from the body when present, falls back to Retry-After in either of its two formats (seconds or an HTTP date), and otherwise uses full jitter. A quota 429 raises at once, because waiting will not fix it.

retry.pyPython
import os
import random
import time
from email.utils import parsedate_to_datetime

import requests

BASE_URL = "https://api.tickerlayer.com"
RETRYABLE = {429, 500, 502, 503, 504}

session = requests.Session()
session.headers["x-api-key"] = os.environ["TICKERLAYER_API_KEY"]


class QuotaExceeded(Exception):
    """The monthly quota is used up. Retrying will not help until it resets."""


def body_of(resp):
    try:
        body = resp.json()
    except ValueError:
        return {}
    return body if isinstance(body, dict) else {}


def server_wait(resp, body):
    """Seconds the server asked us to wait, or None if it did not say."""
    if isinstance(body.get("retryAfterMs"), (int, float)):
        return body["retryAfterMs"] / 1000  # more precise than the header
    value = resp.headers.get("Retry-After")
    if not value:
        return None
    if value.strip().isdigit():
        return float(value)
    try:  # the HTTP-date form: "Wed, 21 Oct 2026 07:28:00 GMT"
        return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
    except (TypeError, ValueError):
        return None


def get(path, params=None, max_attempts=5, base=0.5, cap=20.0):
    """GET a TickerLayer path, retrying 429 and transient 5xx responses."""
    for attempt in range(1, max_attempts + 1):
        resp = session.get(BASE_URL + path, params=params, timeout=10)
        if resp.status_code not in RETRYABLE:
            resp.raise_for_status()  # 400, 401, 403, 404: fix the request instead
            return resp.json()

        body = body_of(resp)
        if body.get("error") == "REST_QUOTA_EXCEEDED":
            raise QuotaExceeded(f"monthly quota of {body.get('limit')} requests used up")
        if attempt == max_attempts:
            resp.raise_for_status()

        hint = server_wait(resp, body)
        if hint is not None:
            delay = hint + random.uniform(0, base)  # honour the server, de-synchronise a little
        else:
            delay = random.uniform(0, min(cap, base * 2 ** attempt))  # full jitter
        print(f"{resp.status_code} on {path} (attempt {attempt}), retrying in {delay:.2f}s")
        time.sleep(delay)


if __name__ == "__main__":
    quote = get("/forex/quote/EURUSD")
    print(quote["symbol"], quote["bid"], quote["ask"], quote["timestamp"])
Output (a normal run, 28 September 2026)
EURUSD 1.137233 1.137242 1790592920001

To exercise the retry path without hammering a real API, we pointed the same function at a local test server that answered 429 (with retryAfterMs), 429 (with only the header) and 503 (with no hint) before succeeding. Each wait follows the rule for its case:

Output (against a local test server)
429 on /forex/quote/EURUSD (attempt 1), retrying in 0.27s
429 on /forex/quote/EURUSD (attempt 2), retrying in 1.32s
503 on /forex/quote/EURUSD (attempt 3), retrying in 1.67s
EURUSD 1.137233 1.137242 1790592920001

JavaScript: the same logic with fetch

Node 18 and later ship fetch, so there is nothing to install. Save the file as retry.mjs so top-level await works.

retry.mjsJavaScript
const BASE_URL = "https://api.tickerlayer.com";
const API_KEY = process.env.TICKERLAYER_API_KEY;
const RETRYABLE = new Set([429, 500, 502, 503, 504]);
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

class QuotaExceeded extends Error {}

// Milliseconds the server asked us to wait, or null if it did not say.
function serverWaitMs(res, body) {
  if (typeof body.retryAfterMs === "number") return body.retryAfterMs;
  const value = res.headers.get("retry-after");
  if (!value) return null;
  if (/^\d+$/.test(value.trim())) return Number(value) * 1000;
  const at = Date.parse(value); // the HTTP-date form
  return Number.isNaN(at) ? null : Math.max(0, at - Date.now());
}

async function get(path, { maxAttempts = 5, baseMs = 500, capMs = 20_000 } = {}) {
  for (let attempt = 1; ; attempt++) {
    const res = await fetch(BASE_URL + path, {
      headers: { "x-api-key": API_KEY },
      signal: AbortSignal.timeout(10_000),
    });
    const body = await res.json().catch(() => ({}));
    if (res.ok) return body;
    if (!RETRYABLE.has(res.status)) {
      throw new Error(`${res.status} ${body.message ?? res.statusText}`); // fix the request
    }
    if (body.error === "REST_QUOTA_EXCEEDED") {
      throw new QuotaExceeded(`monthly quota of ${body.limit} requests used up`);
    }
    if (attempt >= maxAttempts) throw new Error(`${res.status} after ${attempt} attempts`);

    const hint = serverWaitMs(res, body);
    const delay = hint !== null
      ? hint + Math.random() * baseMs                            // honour the server
      : Math.random() * Math.min(capMs, baseMs * 2 ** attempt);  // full jitter
    console.warn(`${res.status} on ${path} (attempt ${attempt}), retrying in ${Math.round(delay)} ms`);
    await sleep(delay);
  }
}

const quote = await get("/forex/quote/EURUSD");
console.log(quote.symbol, quote.bid, quote.ask, quote.timestamp);
Output (a normal run, then the local test server)
EURUSD 1.13725 1.13726 1790592934000

429 on /forex/quote/EURUSD (attempt 1), retrying in 394 ms
429 on /forex/quote/EURUSD (attempt 2), retrying in 1222 ms
503 on /forex/quote/EURUSD (attempt 3), retrying in 1529 ms
EURUSD 1.137233 1.137242 1790592920001

The Python stock prices tutorial builds a full client on the same pattern. Streams need their own version of backoff for reconnects, covered in WebSocket reconnect.

Stop hitting the limit: budget and pace your requests

Retries treat the symptom. Most 429s in market data come from a handful of request patterns, and each has a cheaper shape that returns the same data:

PatternRequestsCheaper shapeRequests
Poll 20 symbols every 2 s36,000 an hourStream them over one WebSocket1 subscribe message
Quote, last trade and previous close per symbol3 per symbolGET /{asset}/snapshot/{symbol}1 per symbol
Open or closed, for 25 markets25GET /markets/status/bulk?items=1
Fundamentals for 100 US stocks100GET /fundamentals/stocks?symbols= (up to 100)1
A year of 1-minute bars for one US stockHundreds of pages at the default 500limit=5000 per pageA tenth as many
Polling through the weekendAll of them wastedRead next_open from /markets/status, then sleep1
Same data, fewer requests. A bulk or batch call counts as one request.

The market status routes, including the bulk form and next_open, are documented in the market status reference. For polling you cannot avoid, work out the budget before you write the loop:

Polling budget calculator

sec
Requests per second
2.00
Requests per day
57,600
Requests per month
1,267,200
Smallest plan that fits
Business
  • Free422×
  • Individual507%
  • Business5%

One request per symbol per poll. A WebSocket subscription replaces all of these requests with one connection. Quotas are listed on pricing; per-second limits are in the X-RateLimit-Limit header.

Compare the monthly total with your plan: 3,000 REST requests a month on the free tier, 250,000 on Individual and 25 million on Business.

Inside the per-second limit, pace rather than burst. X-RateLimit-Limit tells you the budget, so a client can space its requests evenly and never see a 429 at all. This pacer keeps 20% headroom and is safe to share across threads:

pacer.pyPython
import os
import threading
import time

import requests

BASE_URL = "https://api.tickerlayer.com"
session = requests.Session()
session.headers["x-api-key"] = os.environ["TICKERLAYER_API_KEY"]


class Pacer:
    """Spaces requests evenly so a burst never exceeds the per-second budget."""

    def __init__(self, per_second, headroom=0.8):
        self.interval = 1.0 / (per_second * headroom)
        self.next_at = time.monotonic()
        self.lock = threading.Lock()

    def wait(self):
        with self.lock:
            now = time.monotonic()
            start = max(now, self.next_at)
            self.next_at = start + self.interval
        time.sleep(max(0.0, start - now))


first = session.get(f"{BASE_URL}/forex/quote/EURUSD", timeout=10)
first.raise_for_status()
limit = int(first.headers.get("X-RateLimit-Limit", "10"))
pacer = Pacer(limit)
print(f"limit {limit}/s, one request every {pacer.interval * 1000:.0f} ms")

for symbol in ["GBPUSD", "USDJPY", "AUDUSD"]:
    pacer.wait()
    resp = session.get(f"{BASE_URL}/forex/quote/{symbol}", timeout=10)
    resp.raise_for_status()
    quote = resp.json()
    print(symbol, quote["bid"], quote["ask"], "remaining:", resp.headers["X-RateLimit-Remaining"])
Output (example for a key with a 10 per second limit)
limit 10/s, one request every 125 ms
GBPUSD 1.325809 1.325879 remaining: 8
USDJPY 157.0403 157.0573 remaining: 7
AUDUSD 0.701173 0.701233 remaining: 6

Cache what does not change between polls. Symbol lists, market sessions and holiday calendars change rarely, so fetch them once a day. For prices, keep one server-side cache per symbol for as long as your screen can tolerate, and serve every viewer from it instead of calling the API once per page view. When the whole design is a live screen, stop polling: WebSocket vs REST works through when the switch pays off.

Debug a 429 in production

  1. Read the bodyBranch on error. REST_QUOTA_EXCEEDED is a budget problem; REST_RATE_LIMIT_EXCEEDED is a pacing problem.
  2. Log the headersRecord X-RateLimit-Remaining on every response. A value that hits 0 at the same second of every minute points at a scheduled job.
  3. Find the burstLook for cron jobs firing on the minute, retries without jitter, and several services sharing one account.
  4. Check the monthThe dashboard shows usage against the monthly quota; the response headers do not.
  5. Change the shapeBatch, cache, pace or stream, then delete any retry loop you added as a workaround.

Before you ship a client

  • Retries on 429 and transient 5xx only, with a cap on attempts.
  • Retry-After honoured in both the seconds and the HTTP-date form.
  • Full jitter on every computed backoff.
  • Quota 429s raised as an alert, not retried.
  • Requests paced from X-RateLimit-Limit, with headroom.
  • Live prices streamed rather than polled once you watch more than a few symbols.
  • Polling paused while the market is closed.

Rate limits are one part of a healthy integration; symbols, units and timestamps are the others, covered in market data API explained.

Questions

What does 429 Too Many Requests mean?

The server refused your request because you exceeded a rate limit or a quota. The request itself is fine; send it again after the wait the server indicates.

How long should I wait after a 429 error?

As long as the Retry-After header says, either in seconds or until the HTTP date it gives. With no header, back off exponentially with random jitter, starting around half a second.

Is HTTP 429 a client error or a server error?

A client error: it is in the 4xx range because the client sent too many requests. Unlike most 4xx codes, it is safe to retry after waiting.

Do rate-limited requests count against my quota?

On TickerLayer, no. The per-second limit is checked first, so a request refused with REST_RATE_LIMIT_EXCEEDED does not consume monthly quota.

What is the difference between 429 and 503?

A 429 says you are sending too much; a 503 says the server cannot serve requests right now. Both can carry Retry-After, and both deserve backoff with jitter.

How do I fix 429 errors in Python?

Wrap requests in a retry loop that honours Retry-After, backs off with full jitter, caps attempts and does not retry quota errors. Then reduce volume by batching, caching or streaming.

Keep reading

Ready to integrate?

Start with the free tier, explore the docs, and connect via REST or WebSocket in minutes.