Count your own bytes: audit a proxy meter with mitmproxy
Audit a proxy meter yourself: chain mitmproxy in front of the proxy, log request and response bytes, and see which gaps come from TLS, gzip or retries.
To audit a proxy meter yourself, put mitmproxy between your client and the proxy, let a short addon write the size of every request and response to a CSV file, and compare that file with the provider's per-request log. For plain http:// sites the two numbers should agree to the byte. For https:// sites they cannot, and the gap has a shape you can predict.
How to check your proxy provider's meter argued that a usage number you cannot reconstruct is a claim, not a bill. This is the hands-on half: the commands, the numbers we got when we ran them, and a plain list of what those numbers can and cannot prove.
What you need
- Docker and the official image
mitmproxy/mitmproxy:12.1.2, the version we tested. - curl, or the script you actually want to audit.
- A proxy login (
HOST,PORT,USERNAME,PASSWORD) on the HTTP protocol. mitmproxy's upstream mode only speaks HTTP to the next proxy: starting it with--mode upstream:socks5://…stops at once withInvalid server scheme: socks5. Residential ports answer HTTP anyway; on ISP and datacenter services, switch the protocol to HTTP in the dashboard for the test.
How we tested it
We wanted to see both sides of a meter, so we put our own proxy in the provider's seat: Squid 6.13 with a login, logging the request and response bytes of every connection. Behind it sat two local test sites, a copy of httpbin (go-httpbin 2.25.0) and an nginx 1.27 server with gzip switched on. Everything ran in containers on one machine, and no outside provider was involved. When this post says "the meter", it means that Squid log. A real provider's numbers will differ in detail, for the reasons further down.
Step 1: the counting addon
mitmproxy runs Python addons on every request it handles. Save this as count_bytes.py:
import csv
from mitmproxy import ctx, http
from mitmproxy.net.http.http1.assemble import assemble_request_head, assemble_response_head
class CountBytes:
def __init__(self):
self.file = open("bytes.csv", "w", newline="")
self.out = csv.writer(self.file)
self.out.writerow(["host", "path", "status", "sent", "received", "body_on_wire", "body_decoded"])
self.sent = self.received = 0
def write(self, row):
self.out.writerow(row)
self.file.flush()
self.sent += row[3]
self.received += row[4]
def response(self, flow: http.HTTPFlow):
sent = len(assemble_request_head(flow.request)) + len(flow.request.raw_content or b"")
body = len(flow.response.raw_content or b"")
decoded = len(flow.response.get_content(strict=False) or b"")
received = len(assemble_response_head(flow.response)) + body
self.write([flow.request.host, flow.request.path, flow.response.status_code, sent, received, body, decoded])
def error(self, flow: http.HTTPFlow):
self.write([flow.request.host, flow.request.path, "failed", 0, 0, 0, 0])
def done(self):
ctx.log.info(f"sent {self.sent} bytes, received {self.received} bytes")
self.file.close()
addons = [CountBytes()]
received is the status line, the response headers and the body exactly as it crossed the wire, compressed or not. body_decoded is the same body after decompression, kept as a separate column because confusing the two is the most common way a self-audit goes wrong. Requests that never got an answer are written as failed with zeros.
Step 2: chain mitmproxy to the proxy
mkdir -p ~/.mitmproxy
docker run --rm -it -p 127.0.0.1:8080:8080 \
-v "$PWD":/work -w /work \
-v ~/.mitmproxy:/home/mitmproxy/.mitmproxy \
mitmproxy/mitmproxy:12.1.2 \
mitmdump --mode upstream:http://HOST:PORT --upstream-auth USERNAME:PASSWORD \
-s count_bytes.py --set http2=false
--mode upstream: makes mitmproxy forward everything to your proxy instead of connecting to sites itself, and --upstream-auth sends the login, so your client needs no proxy password of its own. --set http2=false keeps every exchange on HTTP/1.1, which makes header sizes plain text you can count. The port is published on 127.0.0.1 only, so nothing else on your network can use it. In our lab we also passed --set ssl_verify_upstream_trusted_ca= with the test sites' self-signed certificates; against real sites, leave that out.
The first start writes mitmproxy's own certificate authority into ~/.mitmproxy. Your client has to trust it, or mitmproxy cannot see inside https:// traffic and the client refuses with curl: (60) SSL certificate problem: unable to get local issuer certificate.
Step 3: send traffic through it
curl -x http://127.0.0.1:8080 --cacert ~/.mitmproxy/mitmproxy-ca-cert.pem \
-o /dev/null https://httpbin.org/bytes/1000
For a Python script built on requests, two environment variables do the same job, and we confirmed it with requests 2.32.5:
HTTPS_PROXY=http://127.0.0.1:8080 \
REQUESTS_CA_BUNDLE=~/.mitmproxy/mitmproxy-ca-cert.pem \
python3 your_script.py
Stop mitmdump with Ctrl+C and the addon prints the totals, sent 543 bytes, received 13729 bytes in our run, with bytes.csv beside it. This is the file from that run. In the lab, origin was the httpbin copy and shop the nginx server:
host,path,status,sent,received,body_on_wire,body_decoded
origin,/bytes/1000,200,84,1190,1000,1000
shop,/catalogue.html,200,123,11721,11470,182045
origin,/status/429,429,84,203,0,0
origin,/status/503,503,84,205,0,0
origin,/status/503,503,84,205,0,0
origin,/status/503,503,84,205,0,0
nowhere.invalid,/,failed,0,0,0,0
If the upstream login is wrong, mitmdump logs Upstream proxy HOST:PORT refused HTTP CONNECT request: 407 Proxy Authentication Required and curl gets a 502 from mitmproxy. Fixing a 407 covers the usual causes.
Step 4: line it up with the provider's log
Export the provider's per-request log for the same minutes, match rows by time and host, and set your received column next to their byte column. Run one request per connection while you do this: separate curl commands each open their own, which keeps one of your rows facing one of theirs.
What the numbers can prove
Plain http: byte for byte
For http://plain:8080/bytes/1000, our CSV said 179 bytes sent and 1,294 received, and the meter logged 179 and 1,294. A proxy that reads plain HTTP sees exactly what you see, including the headers it adds itself (our Squid added Via and Cache-Status, which reached mitmproxy as well). Any real gap here is worth an email.
https: the meter counts a tunnel
For an https:// site, the client asks the proxy to open a CONNECT tunnel and encrypts everything inside it. A meter at the proxy can only count what crosses that tunnel: the TLS handshake, the site's certificate and a few bytes of framing on every encrypted record, on top of the headers and body. mitmproxy decrypts, so it counts only the HTTP inside. Here is what that looked like:
| What we fetched | mitmproxy received |
Meter, bytes back through the tunnel |
|---|---|---|
| One 1,000-byte response, new connection | 1,190 | 3,858 |
| Five of them on one kept-alive connection | 5,950 | 8,706 |
| Ten 100,000-byte responses, one connection | 1,001,920 | 1,006,216 |
| The same ten, a new connection each | 1,001,920 | 1,031,020 |
The handshake cost about 2.7 KB per connection, and our test sites sent one small self-signed certificate. A real site sends a chain, often several kilobytes, so expect more. On tiny responses over fresh connections the handshake is most of what gets counted; on large responses over a reused connection the gap fell to 0.4%, and with a new connection per 100 KB response it was 2.9%. Notice too that five requests made one tunnel row: if a log line stands for a connection, it may hold several of your pages.
Compression: count what crossed the wire
The test catalogue page was 182,045 bytes of HTML and 11,470 bytes as gzip. Python requests asks for gzip by default and hands you the decompressed body, so len(r.content) printed 182,045 for a transfer the meter saw as 13,732 bytes. Count body_on_wire, or set Accept-Encoding: identity while you measure.
Retries are traffic too
curl --retry 2 against a 503 sent three requests, three rows in the CSV, and all three went through one tunnel that the meter logged as 3,305 bytes. The site answered every attempt, so every attempt cost something. Retries without burning bandwidth has the patterns that avoid paying for the same refusal twice.
A site's error is not a failed connection
The 429 came back as 203 bytes of HTTP inside 2,849 bytes of tunnel: the site answered, so it is traffic. A host that does not exist produced a failed row with zeros, and the meter logged the attempt with nothing received from any site.
What the numbers cannot prove
- mitmproxy changes the connection it measures. It opens its own TLS session to the site, so the handshake the meter counts is mitmproxy's, not your client's, and the site sees mitmproxy's TLS fingerprint. Audit against a neutral endpoint such as
httpbin.org, not the site you scrape. - HTTP/2 headers are smaller. We switched HTTP/2 off so headers are counted as HTTP/1.1 text. Over HTTP/2 they are compressed on the wire, and the addon would overstate them.
- It sees only what you route through it. Another program using the same proxy login shows up in the provider's log and not in your file.
- It cannot see inside the provider. A local count bounds what you should expect. Large gaps on plain http, or on big transfers over kept-alive https connections, are the ones worth raising.
Charles or Fiddler instead of a script
If you prefer a window to a terminal, Charles and Fiddler Everywhere both list every request with its size and can forward to another proxy. Charles calls it External Proxies, with separate entries for HTTP, HTTPS and SOCKS and support for a Basic login. Fiddler Everywhere puts it under Gateway, Manual proxy configuration; its documentation does not mention a login for that upstream, so pair it with an IP allowlist. The comparisons above apply unchanged, and so does the certificate step: both need their own root certificate trusted before they can show https:// traffic.
Quick answers
Can I check a proxy meter without the provider's help? You can measure your side. Comparing it needs their per-request log, which is why it is worth asking for one before you buy.
Why is the provider's number higher than mine for https? Their meter counts the encrypted tunnel, handshake and framing included, and mitmproxy counts the HTTP inside it. The gap is largest for small responses on fresh connections.
Does a 429 or a block page count as traffic? On any meter that counts bytes from the site, yes: the site sent them.
How ProxyPanda's meter is written down
Our honesty page states it: we bill response body plus response headers, measured at our proxy, and request bytes are free. Inside an https tunnel our proxy cannot see the site's status code, so the meter counts the bytes that crossed it. Timeouts, resets and errors from our own gateway bill zero and stay on the log marked free, and the whole log exports as CSV. If your count and ours disagree by more than the tolerance written there, send both and we re-bill at yours while we look into it. The metering explainer has the questions worth asking any provider, and pricing has the rates the bytes are multiplied by.
Next step
Start mitmdump as above and fetch https://httpbin.org/bytes/100000 ten times in one curl command, which reuses one connection, then ten times in separate commands. Put both totals next to your provider's log for those minutes: the first pair should sit close together, and the second shows you what a handshake costs on that network. If the numbers are hard to read, bring the CSV to Discord with the login removed.