Explainer29 September 20267 min read

How per-GB proxy billing works, and what a scrape will cost

A formula and a short script for estimating how many gigabytes a scraping job will use on a per-GB proxy, before you order any bandwidth.

"How many gigabytes do I need?" is the question we answer most often in Discord, and the honest reply is that it depends on your pages. The good part is that you can measure the thing it depends on in about ten minutes, and turn it into an estimate you can trust within a reasonable margin.

This post is about per-GB billing, which is how our residential line is sold. ISP and datacenter are priced by the IP and the term; the pricing page shows what traffic an IP plan includes.

What a per-GB meter counts

A proxy sits between you and the site, so it sees every byte the site sends back to you. Providers differ on the details, and the details are where bills go wrong:

  • Which direction. Some bill what you send as well as what you receive. We bill the response only: body plus headers, measured at our proxy.
  • Failed requests. A timeout or a connection reset returns nothing. With us, those and errors from our own gateway, such as a 407 or a 502, bill zero and still show on your log marked free. A 429 or a 5xx page from the site is a response, and its bytes count.
  • Blocks with a body. A 403 page or a captcha arrives with HTML, so it is traffic like any other response, here and everywhere else.

The full rules are under metering. Whoever you buy from, find their version of that list before estimating anything.

The formula

gigabytes ≈ pages × average page weight × (1 + retry share) ÷ 1,000,000,000
cost      ≈ gigabytes × your rate per GB
  • Pages is how many responses the job needs, counting list pages, detail pages and anything you fetch twice.
  • Average page weight is bytes on the wire for one response, headers included. You measure this below.
  • Retry share is the fraction of extra responses that come back with a body you did not want, such as blocks and captchas. Take it from a previous run, or start with a cautious guess and correct it after the first batch.
  • Your rate per GB comes from the pricing page. It falls as the order gets bigger, so read the rung that matches the quantity you plan to buy.

Measuring page weight

Take twenty or so URLs that look like the job: a mix of the page types you will fetch, in the proportions you will fetch them. Then weigh them through the proxy, reading the compressed bytes before your client unpacks them.

import requests

PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
HEADERS = {"Accept-Encoding": "gzip, deflate"}


def wire_bytes(url):
    r = requests.get(url, proxies=PROXIES, headers=HEADERS, stream=True, timeout=30)
    body = r.raw.read(decode_content=False)
    head = sum(len(k) + len(v) + 4 for k, v in r.headers.items())
    return len(body) + head


sample = [line.strip() for line in open("sample-urls.txt") if line.strip()]
weights = [wire_bytes(url) for url in sample]
average = sum(weights) / len(weights)
print(f"average response on the wire: {average / 1000:.1f} KB")

pages = int(input("How many responses will the job need? "))
retry_share = float(input("Extra unwanted responses, as a fraction (for example 0.05): "))
gigabytes = pages * average * (1 + retry_share) / 1_000_000_000
print(f"estimated traffic: {gigabytes:.2f} GB")

rate = float(input("Your rate per GB from the pricing page: "))
print(f"estimated cost: {gigabytes * rate:.2f}")

HOST, PORT, USERNAME and PASSWORD are on the service's page in our dashboard. The sample costs a little bandwidth, which is the point: you are paying a few pages to avoid guessing about a few hundred thousand.

What moves the number most

Compression. Most sites send HTML gzipped when asked, often at a fraction of its unpacked size. Most HTTP libraries ask by default. If you have turned it off, or your tool strips the header, you may be paying for several times the bytes you need.

A browser instead of a request. A headless browser downloads the scripts, fonts, images and tracking calls the page asks for. That can make one page weigh many times what its HTML alone does. If you only need the HTML, fetch it directly, or block the resource types you do not need.

The endpoint you pick. Many sites load their data from a JSON endpoint you can call yourself. A JSON response is usually a small fraction of the page that renders it.

Fetching things twice. Pagination loops that revisit pages, and retries after soft blocks, both inflate the page count. Retrying without burning bandwidth covers the second.

Check the estimate against the meter

Run the first real batch small, then compare three numbers: your estimate, your own byte count, and the dashboard. If your count and the dashboard agree and the estimate was off, fix the inputs and estimate again. If your count and the dashboard disagree, that is a different conversation, and how to check your provider's meter is the place to start it.

Not sure about your inputs? Post the page types and rough counts in Discord and we will help you size it, including telling you when a smaller order is enough.

Got a follow-up question?

Ask it in Discord. The answer helps whoever reads the thread next.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord