How to fix 429 Too Many Requests when scraping
429 Too Many Requests means the website is rate limiting you. Read Retry-After, work out which limit you hit, and pace requests with a Python limiter.
A 429 Too Many Requests error means the website has counted your requests and decided you are over its limit for now. The fix is to slow down: read the Retry-After header, pause every request to that host for as long as it says, then resume at a lower, steady rate. Rotating IPs helps only when the limit is per IP.
Below: where the status comes from, how to read Retry-After in both formats, the kinds of rate limit, and a per-host rate limiter in Python that honours the header and knows when to give up. Every snippet uses HOST, PORT, USERNAME and PASSWORD as placeholders for the values on your service's page in the dashboard.
What does 429 Too Many Requests mean?
It is the website's answer. The proxy's own refusals look different: a 407 for a bad login, or a 502 when it cannot reach the site. A 429 means your request arrived, the site counted how many it had seen from you recently, and said no.
The site decides what "you" means: your IP address, your account, your API key, a cookie, or a mix. That choice decides which fixes work. The rate limit entry in the glossary has the one-line version.
To confirm the 429 comes from the target, send the same few requests through the same proxy to https://api.ipify.org. If those come back clean while the target keeps answering 429, the target is limiting you.
Read the Retry-After header first
Many sites tell you how long to wait. Look at the headers of one 429 response:
curl -s -o /dev/null -D - -x "http://USERNAME:PASSWORD@HOST:PORT" "https://TARGET_SITE/page/1"
A rate-limited answer looks something like this:
HTTP/2 429
content-type: text/html; charset=utf-8
retry-after: 30
Retry-After comes in two formats, and your code needs to handle both:
- A number of seconds, such as
Retry-After: 30. - An HTTP date, such as
Retry-After: Wed, 21 Oct 2015 07:28:00 GMT, meaning "not before this moment".
Python's standard library can parse the date form, so this needs no extra packages:
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
def retry_after_seconds(value):
if not value:
return None
value = value.strip()
if value.isdigit():
return int(value)
try:
when = parsedate_to_datetime(value)
except (TypeError, ValueError):
return None
if when.tzinfo is None:
when = when.replace(tzinfo=timezone.utc)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
print(retry_after_seconds("120"))
print(retry_after_seconds("Wed, 21 Oct 2015 07:28:00 GMT"))
print(retry_after_seconds("soon"))
That prints 120, then 0.0 because the date is in the past, then None for a value it cannot read. None means "no instruction", and your own backoff takes over.
Some APIs also send X-RateLimit-Remaining and X-RateLimit-Reset, which warn you before you hit the limit.
Which rate limit did you hit?
| Limit counted per | Typical sign | Does a new IP help? |
|---|---|---|
| IP address | 429s start after a burst and clear after a pause; a fresh IP works at once | Yes |
| Account or API key | 429s follow you across IPs while you are signed in or sending the key | No |
| Session or cookie | A fresh session clears it; the same cookie jar keeps failing on any IP | A new session may clear it, but resetting sessions to dodge a limit is evasion; slow down instead |
| Route or whole site | Heavy pages such as search return 429 while light pages still load | No |
The cheap test: after a 429, send one request through a different IP with no cookies. If it succeeds, the limit is at least partly per IP. A limit on your account or API key is part of the deal you signed up for; stay under it, or ask the site for a higher quota.
Why rotating proxies fix only some 429 errors
A per-IP limit is the one case where rotation changes the arithmetic. With residential on Randomize IP, each new connection leaves from a different address, so no single address builds up a count. With a list of static datacenter or ISP IPs, you share the load across them yourself; rotating proxies in Python shows both setups, including the kept-alive session that quietly stops rotation.
Rotation does not make your total rate polite. A thousand requests a minute from a thousand addresses is still a pattern a site can see. Use rotation to spread a reasonable rate.
A per-host rate limiter in Python
The pattern that works is a shared limiter per host: every worker asks it for permission before sending, and when any worker receives a 429, the whole host pauses. Per-request retries on their own are not enough, because ten threads each retrying after their own backoff keep hitting the site in the middle of its cool-down.
import random
import threading
import time
from concurrent.futures import ThreadPoolExecutor
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.parse import urlsplit
import requests
PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
START_RATE = 1.0
MIN_RATE = 0.1
MAX_WAIT = 600
MAX_ATTEMPTS = 5
stop = threading.Event()
class GiveUp(Exception):
pass
def retry_after_seconds(value):
if not value:
return None
value = value.strip()
if value.isdigit():
return int(value)
try:
when = parsedate_to_datetime(value)
except (TypeError, ValueError):
return None
if when.tzinfo is None:
when = when.replace(tzinfo=timezone.utc)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
class HostLimiter:
def __init__(self, rate):
self.start_rate = rate
self.lock = threading.Lock()
self.rate = {}
self.next_slot = {}
self.pauses = {}
def wait(self, host):
while True:
with self.lock:
now = time.monotonic()
rate = self.rate.setdefault(host, self.start_rate)
slot = max(now, self.next_slot.get(host, now))
self.next_slot[host] = slot + 1 / rate
pauses = self.pauses.get(host, 0)
time.sleep(slot - now)
with self.lock:
if self.pauses.get(host, 0) == pauses:
return
def back_off(self, host, seconds):
with self.lock:
self.rate[host] = max(MIN_RATE, self.rate[host] / 2)
self.next_slot[host] = time.monotonic() + seconds
self.pauses[host] = self.pauses.get(host, 0) + 1
def recover(self, host):
with self.lock:
self.rate[host] = min(self.start_rate, self.rate[host] * 1.05)
def fetch(url, limiter):
host = urlsplit(url).hostname
for attempt in range(MAX_ATTEMPTS):
limiter.wait(host)
if stop.is_set():
return None
response = requests.get(url, proxies=PROXIES, timeout=(5, 30))
if response.status_code != 429:
limiter.recover(host)
return response
wait = retry_after_seconds(response.headers.get("Retry-After"))
if wait is None:
wait = min(300, 10 * 2 ** attempt) * random.uniform(1, 1.5)
if wait > MAX_WAIT:
raise GiveUp(f"{host} asked for a {wait:.0f}s wait")
print(f"429 from {host}: pausing the host for {wait:.1f}s")
limiter.back_off(host, wait)
raise GiveUp(f"{host} still answers 429 after {MAX_ATTEMPTS} attempts")
def crawl(urls, workers=4):
limiter = HostLimiter(START_RATE)
def job(url):
if stop.is_set():
return url, None
try:
return url, fetch(url, limiter)
except GiveUp as reason:
if not stop.is_set():
stop.set()
print(f"Stopping the run: {reason}")
return url, None
with ThreadPoolExecutor(workers) as pool:
for url, response in pool.map(job, urls):
if response is not None:
print(response.status_code, url)
if __name__ == "__main__":
crawl([f"https://TARGET_SITE/page/{n}" for n in range(1, 21)])
How the limiter behaves
- Evenly spaced slots.
wait()hands each request the next free time slot for its host,1 / rateseconds after the previous one. It is a token bucket that holds a single token, so there are no bursts. AtSTART_RATE = 1.0the host sees one request a second, however many workers you run. - Workers are not the rate. The thread count decides how many slow responses can be waiting at once; the limiter decides the pace.
- One 429 pauses the whole host.
back_off()pushes the next slot past theRetry-Aftertime and halves the rate. Threads that were already waiting for a slot notice the pause and queue again behind it. - Slow recovery. Each success raises the rate by 5%, never above where you started.
- A fallback when there is no header. Without
Retry-After, the wait grows from about 10 seconds to a few minutes, with some randomness so separate jobs do not return in step.
We ran it against a local test server limited to five requests a second, through a password-protected test proxy. At a starting rate of four a second, 30 pages came back with no 429 at all. At fifteen a second, it met three 429s, in both header formats, paused, halved its pace and finished all 40 pages.
When should a scraper stop?
The limiter gives up in two cases, and both are deliberate:
- The site asks for a long wait. A
Retry-Afterof an hour is the site telling you to come back later.MAX_WAITturns that into a clean stop, so a scheduled job can resume at its next run. - The 429s keep coming. Five in a row on the same host, even after backing off, means your idea of the limit is wrong. Stopping costs you nothing; carrying on sends hundreds more requests, and a site that sees you ignore its 429s may move you on to 403s and block pages.
After a stop, lower START_RATE before the next run. A good first guess for a site you do not know is one request a second per host. Proxy concurrency shows how to measure a safe level for a particular target.
429 vs 403: what is the difference?
A 429 says "too many, wait". It is temporary and often tells you for how long. A 403 says "not you" and does not end by itself: your IP, fingerprint or behaviour has been judged, and waiting a minute changes nothing. The fixes for 403 are different, and why your scraper started getting blocked goes through them in order. Some sites send 403 or a captcha page where others would send 429, so if a 403 clears after a pause, treat it as a rate limit.
What does a 429 cost on a per-GB proxy?
On our residential meter, a 429 is a response from the site, so its bytes are traffic like any other. For an HTTPS site the status code travels inside the encrypted tunnel, where our proxy cannot read it, and the meter counts the bytes that crossed. A 429 is usually short, a small fraction of a real page, so one costs little. A loop that ignores Retry-After and collects hundreds of them adds up. Only connection failures (timeouts, resets, and errors from our own gateway such as a 407 or a 502) are billed at zero; the rules are under metering.
The bigger waste lies elsewhere: each ignored 429 lowers your standing with the site, and the block page that often follows is heavier, and metered the same way. Retrying without burning bandwidth covers the other failure types.
Quick answers
Is a 429 error my proxy's fault? Rarely. It is the website's reply. Check by sending the same requests through the proxy to an IP echo service: if those work, the target is limiting you.
How long should I wait after a 429?
As long as Retry-After says. Without it, start around 10 seconds and double on each repeat, up to a few minutes.
Will more proxies fix 429 Too Many Requests? Only if the limit is per IP. Limits on an account, API key or route follow you to every IP.
Why do I get 429 even with rotating proxies? Your requests may share a session or cookie, the site may limit the route regardless of IP, or a kept-alive connection may be sending everything from one exit.
Where to go from here
Start with the limiter at one request a second on a small batch, and raise it only while the 429 count stays at zero. If you want to try this through us, a small top-up is enough, and the pricing page has the rates. If a target keeps limiting you at a gentle pace, describe it in Discord with the headers from one 429 response.