Tutorial29 September 20267 min read

Handling blocks and retries without burning proxy bandwidth

On a per-GB proxy, a careless retry loop pays for the same block page again and again. Which failures to retry, how to back off, and when to stop early.

A retry loop that looks harmless on your laptop can quietly become the biggest line on a per-GB bill. The usual culprit is a site that has started answering with a block page: every attempt downloads the same few hundred kilobytes of "please verify you are human", and the loop keeps asking.

This post is about spending bandwidth only on responses you want. It assumes a per-GB product such as our residential line; on per-IP lines the same habits save time and keep your IPs in better standing.

What a failed attempt costs

It depends on what came back.

Outcome Body from the site? On our meter
Timeout, connection reset, an error from our gateway such as a 407 or 502 No Free, shown on the log
429, 5xx from the site Usually small Metered like any response
403 page, captcha, "verify you are human" Yes Metered like any response
200 carrying a block page Yes Metered like any response

Only the first row is free: the connection failed, and nothing from the site crossed the proxy. Every other row is the site answering, and those bytes are traffic. For an HTTPS site the status code travels inside the encrypted tunnel, where our proxy cannot read it, so the meter counts the bytes that crossed and cannot tell a 429 from a 200. The rules are under metering.

A 429 or a 503 is usually a short response, a small fraction of a real page, so retrying one after a proper wait costs little. Block pages are heavier, and a loop that fetches them again and again is where retries burn money.

Retry what can change

Sort responses before deciding to try again:

  • Retry with backoff: timeouts, connection errors, 429, 502, 503, 504. These are transient, or the site asking you to wait. Each 429 or 5xx you collect is a small metered response, so wait properly between attempts and cap how many you make.
  • Retry once on a fresh IP, then stop: 403 and obvious block pages. With rotation, one more attempt on a new address is reasonable. A second block means the problem is your request, your rate or your fingerprint, and why your scraper started getting blocked is the checklist.
  • Do not retry: 404, 410, 400. The answer will not change.

Back off, and honour Retry-After

Exponential backoff with some randomness keeps a pool of workers from retrying in lockstep. When the site sends Retry-After, it has told you the number to use.

import random
import time


def backoff(attempt):
    return min(60, 2 ** attempt) * random.uniform(0.5, 1.0)


def retry_after(response):
    value = response.headers.get("Retry-After", "")
    return int(value) if value.isdigit() else None

Stop the download early

With stream=True, requests reads the status line and headers first and fetches the body only when you ask for it. That gives you two cheap exits: refuse a status you do not want before any body arrives, and read a first chunk to spot a block page before downloading the rest. Closing the response early means most of an unwanted body never crosses the proxy.

import requests

PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
RETRYABLE = {429, 502, 503, 504}
BLOCK_MARKERS = (b"captcha", b"verify you are human", b"access denied")
MAX_BYTES = 3_000_000


def fetch(url, attempts=4):
    blocked_once = False
    for attempt in range(attempts):
        try:
            response = requests.get(url, proxies=PROXIES, timeout=30, stream=True)
        except (requests.ConnectionError, requests.Timeout):
            time.sleep(backoff(attempt))
            continue

        with response:
            if response.status_code in RETRYABLE:
                time.sleep(retry_after(response) or backoff(attempt))
                continue
            if response.status_code == 403:
                if blocked_once:
                    return None
                blocked_once = True
                continue
            if response.status_code != 200:
                return None
            if int(response.headers.get("Content-Length") or 0) > MAX_BYTES:
                return None

            chunks = response.iter_content(65_536)
            first = next(chunks, b"")
            if any(marker in first.lower() for marker in BLOCK_MARKERS):
                if blocked_once:
                    return None
                blocked_once = True
                continue
            return first + b"".join(chunks)
    return None

Choose BLOCK_MARKERS from the block pages your target serves; open one in a browser and copy a phrase from it. HOST, PORT, USERNAME and PASSWORD come from the service's page in our dashboard.

Pause when the whole host is saying no

Per-URL retries do not notice when every URL on a host is failing. Keep a short window of recent outcomes per host, and when most of them are blocks, stop sending for a while.

from collections import defaultdict, deque

recent = defaultdict(lambda: deque(maxlen=20))


def record(host, blocked):
    recent[host].append(blocked)


def should_pause(host):
    window = recent[host]
    return len(window) == window.maxlen and sum(window) > window.maxlen // 2

A pause costs nothing. A hundred workers downloading block pages through the same stretch is a real line on the bill.

Ask for fewer bytes in the first place

  • Keep compression on. Most HTTP clients send Accept-Encoding: gzip by default. Check that yours does.
  • Use conditional requests for pages you revisit. Send the ETag you saw last time as If-None-Match; an unchanged page comes back as a 304 with no body.
  • Prefer the JSON the page loads. A site's own data endpoint is usually a fraction of the rendered page.
  • In a headless browser, block what you do not need. Images, fonts and media often outweigh the HTML many times over.
etags = {}


def fetch_if_changed(url):
    headers = {"If-None-Match": etags[url]} if url in etags else {}
    response = requests.get(url, headers=headers, proxies=PROXIES, timeout=30)
    if response.status_code == 304:
        return None
    if "ETag" in response.headers:
        etags[url] = response.headers["ETag"]
    return response.content

To see what all this saves you, compare your own byte count with the meter before and after; how per-GB billing works has the script. Questions about a specific target are welcome in Discord.

Got a follow-up question?

Ask it in Discord. The answer helps whoever reads the thread next.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord