Why your scraper started getting blocked
It worked for three weeks and now everything is a 403. Usually it is not the proxies. Here are the five causes, in the order they are worth checking.
Nothing changed on your side, and now every request comes back 403. The instinct is to blame the proxies and buy better ones. Before you spend anything, work through this list. It is roughly ordered by how often each turns out to be the culprit.
1. Your rate went up without you noticing
The most common cause, by a distance. You added concurrency, or the target's pages got slower so your worker pool started overlapping, or a retry loop is quietly firing three times per URL.
Measure your requests per minute against the target, instead of trusting the number you set. If it crept up, drop it and see if the blocks stop.
2. Your fingerprint gives you away
A residential IP does not help if the request arriving from it is obviously automated. Sites increasingly check:
- TLS fingerprint (JA3/JA4). Python's
requestshas a distinctive TLS handshake. A real Chrome does not look like that. Libraries likecurl_cffiexist to fix exactly this. - Header order and casing. Real browsers send a specific set in a specific order. Most HTTP libraries do not.
- Missing headers. No
Accept-Language, noSec-Ch-Ua, noRefereron a page you supposedly clicked to.
If your fingerprint is the problem, upgrading from datacenter to residential will not help. You will just pay more per block.
3. The site deployed something
Sites change. A new WAF, a new bot-management vendor, or a tightened rule set on the existing one. You will often see the change as a sudden clean break: fine on Tuesday, 403 on Wednesday, nothing in between.
Test with a real browser from a residential IP by hand. If a normal browser also gets challenged, it is them, not you.
4. You are reusing sticky sessions too long
If you hold one IP for a long session and push a thousand requests through it, that IP now looks like a data centre regardless of where it lives. Sticky sessions are for keeping a login alive, not for volume.
Shorten the session, or rotate per request for the parts that do not need continuity.
5. The IPs are burnt
It happens, and it is worth checking properly before assuming it. First confirm the IP is the type you paid for: checking what kind of IP you were sold shows how. Then a quick test: run the same request through a fresh IP from a different provider, or through your own home connection. If your connection sails through and the proxy does not, the pool has a problem on that target.
Tell your provider. A decent one wants to know, because burnt IPs cost them more than they cost you.
The order to spend money in
- Slow down. Free.
- Fix your fingerprint. Free, a few hours of work.
- Shorten sticky sessions. Free.
- Move datacenter to residential. It costs more per job, because residential is billed by the gigabyte.
Most people jump straight to step four, and a good share of them were on step one all along.
While you are testing
On our per-GB meter, a request that brings nothing back from the site (a timeout, a reset, an error from our own gateway) bills zero and still shows on your log marked free. Anything the site sends back is different: a 429, a 5xx, a 403 or a captcha page arrives with a body, and those bytes are traffic like any other response. So a diagnostic run that hits the wall is cheap, but it is not free, and the cheapest diagnostic is a small one. Retrying without burning bandwidth shows how to stop early.
The exact rules for what is metered are on the honesty page.
Still stuck after all five? Bring it to Discord with the status codes and a sample response. Someone has usually met the same wall.