Handling blocks and retries without burning proxy bandwidth
On a per-GB proxy, a careless retry loop pays for the same block page again and again. Which failures to retry, how to back off, and when to stop early.
A retry loop that looks harmless on your laptop can quietly become the biggest line on a per-GB bill. The usual culprit is a site that has started answering with a block page: every attempt downloads the same few hundred kilobytes of "please verify you are human", and the loop keeps asking.
This post is about spending bandwidth only on responses you want. It assumes a per-GB product such as our residential line; on per-IP lines the same habits save time and keep your IPs in better standing.
What a failed attempt costs
It depends on what came back.
| Outcome | Body from the site? | On our meter |
|---|---|---|
| Timeout, connection reset, an error from our gateway such as a 407 or 502 | No | Free, shown on the log |
| 429, 5xx from the site | Usually small | Metered like any response |
| 403 page, captcha, "verify you are human" | Yes | Metered like any response |
| 200 carrying a block page | Yes | Metered like any response |
Only the first row is free: the connection failed, and nothing from the site crossed the proxy. Every other row is the site answering, and those bytes are traffic. For an HTTPS site the status code travels inside the encrypted tunnel, where our proxy cannot read it, so the meter counts the bytes that crossed and cannot tell a 429 from a 200. The rules are under metering.
A 429 or a 503 is usually a short response, a small fraction of a real page, so retrying one after a proper wait costs little. Block pages are heavier, and a loop that fetches them again and again is where retries burn money.
Retry what can change
Sort responses before deciding to try again:
- Retry with backoff: timeouts, connection errors, 429, 502, 503, 504. These are transient, or the site asking you to wait. Each 429 or 5xx you collect is a small metered response, so wait properly between attempts and cap how many you make.
- Retry once on a fresh IP, then stop: 403 and obvious block pages. With rotation, one more attempt on a new address is reasonable. A second block means the problem is your request, your rate or your fingerprint, and why your scraper started getting blocked is the checklist.
- Do not retry: 404, 410, 400. The answer will not change.
Back off, and honour Retry-After
Exponential backoff with some randomness keeps a pool of workers from retrying in lockstep. When the site sends Retry-After, it has told you the number to use.
import random
import time
def backoff(attempt):
return min(60, 2 ** attempt) * random.uniform(0.5, 1.0)
def retry_after(response):
value = response.headers.get("Retry-After", "")
return int(value) if value.isdigit() else None
Stop the download early
With stream=True, requests reads the status line and headers first and fetches the body only when you ask for it. That gives you two cheap exits: refuse a status you do not want before any body arrives, and read a first chunk to spot a block page before downloading the rest. Closing the response early means most of an unwanted body never crosses the proxy.
import requests
PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
RETRYABLE = {429, 502, 503, 504}
BLOCK_MARKERS = (b"captcha", b"verify you are human", b"access denied")
MAX_BYTES = 3_000_000
def fetch(url, attempts=4):
blocked_once = False
for attempt in range(attempts):
try:
response = requests.get(url, proxies=PROXIES, timeout=30, stream=True)
except (requests.ConnectionError, requests.Timeout):
time.sleep(backoff(attempt))
continue
with response:
if response.status_code in RETRYABLE:
time.sleep(retry_after(response) or backoff(attempt))
continue
if response.status_code == 403:
if blocked_once:
return None
blocked_once = True
continue
if response.status_code != 200:
return None
if int(response.headers.get("Content-Length") or 0) > MAX_BYTES:
return None
chunks = response.iter_content(65_536)
first = next(chunks, b"")
if any(marker in first.lower() for marker in BLOCK_MARKERS):
if blocked_once:
return None
blocked_once = True
continue
return first + b"".join(chunks)
return None
Choose BLOCK_MARKERS from the block pages your target serves; open one in a browser and copy a phrase from it. HOST, PORT, USERNAME and PASSWORD come from the service's page in our dashboard.
Pause when the whole host is saying no
Per-URL retries do not notice when every URL on a host is failing. Keep a short window of recent outcomes per host, and when most of them are blocks, stop sending for a while.
from collections import defaultdict, deque
recent = defaultdict(lambda: deque(maxlen=20))
def record(host, blocked):
recent[host].append(blocked)
def should_pause(host):
window = recent[host]
return len(window) == window.maxlen and sum(window) > window.maxlen // 2
A pause costs nothing. A hundred workers downloading block pages through the same stretch is a real line on the bill.
Ask for fewer bytes in the first place
- Keep compression on. Most HTTP clients send
Accept-Encoding: gzipby default. Check that yours does. - Use conditional requests for pages you revisit. Send the
ETagyou saw last time asIf-None-Match; an unchanged page comes back as a304with no body. - Prefer the JSON the page loads. A site's own data endpoint is usually a fraction of the rendered page.
- In a headless browser, block what you do not need. Images, fonts and media often outweigh the HTML many times over.
etags = {}
def fetch_if_changed(url):
headers = {"If-None-Match": etags[url]} if url in etags else {}
response = requests.get(url, headers=headers, proxies=PROXIES, timeout=30)
if response.status_code == 304:
return None
if "ETag" in response.headers:
etags[url] = response.headers["ETag"]
return response.content
To see what all this saves you, compare your own byte count with the meter before and after; how per-GB billing works has the script. Questions about a specific target are welcome in Discord.