Setup guide
Scrapy proxy setup with request.meta
Scrapy ships with a proxy middleware that is already switched on. You give a request a proxy URL in its meta, and Scrapy builds the authentication header for you.
Before you start
- Scrapy 2.13 or newer, which added the async start() method used below. On an older version, rename it to start_requests and drop the async keyword.
- The HOST, PORT, USERNAME and PASSWORD from your service page.
- A Scrapy project, or a single spider file you run with scrapy runspider.
Your connection details
Sign in to the dashboard and open the service you bought. Its page shows the host, port, username and password for that service. There is no single ProxyPanda address to remember, so copy them from there each time. The code in these guides uses the placeholders below; replace each one with your own value.
- HOST
- The proxy address for this service, exactly as the service page shows it.
- PORT
- The port to connect to. Copy it with the host, since it can differ from one service to the next.
- USERNAME
- Your proxy login. It is separate from the email you use for the dashboard.
- PASSWORD
- Copy it in full. On residential, the options you choose in the dashboard are added to the end of the password, so a password typed from memory loses them.
You can skip the username and password by adding the IP address you connect from to the service’s allowlist. Then only HOST and PORT go into your code.
The examples use HTTP, which reaches https sites through an encrypted tunnel. SOCKS5 works too: residential ports accept both, and ISP and datacenter proxies switch protocol in the dashboard.
Steps
Set meta["proxy"] on the request
HttpProxyMiddleware reads the proxy key from each request. When the URL carries USERNAME:PASSWORD, the middleware moves them into a Proxy-Authorization header, so there is nothing else to configure.
ip_spider.py import scrapy PROXY = "http://USERNAME:PASSWORD@HOST:PORT" class IpSpider(scrapy.Spider): name = "ip" async def start(self): yield scrapy.Request("https://httpbin.org/ip", meta={"proxy": PROXY}) def parse(self, response): yield response.json()Run the spider and read the exit IP
httpbin.org/ip replies with the address it saw. The origin field in out.json should hold the proxy IP.
terminal scrapy runspider ip_spider.py -O out.json cat out.jsonCover every request from the environment
To proxy a whole crawl without touching each Request, export https_proxy before starting Scrapy. The middleware reads it when the crawl starts. A proxy set in meta still wins over the environment for that one request.
terminal export https_proxy="http://USERNAME:PASSWORD@HOST:PORT" scrapy crawl your_spider
Rotating and sticky IPs
Scrapy keeps many requests in flight at once. On residential with Randomize IP they leave from many addresses, which spreads a crawl of listing and search pages across the pool.
Logins, carts and anything with a session cookie need one address from start to finish. Use Sticky IP for those spiders, or an ISP IP that stays the same for its whole term.
Common errors and fixes
TunnelError: Could not open CONNECT tunnel with proxy ... {'status': 407}
The credentials in the proxy URL were refused. Copy USERNAME and PASSWORD again from the service page and percent-encode characters such as @ or : inside them.
Crawled (400) from an IP-echo service while the proxy connected fine
Some IP-echo services turn away Scrapy’s default client even without a proxy. Test against httpbin.org/ip, as above, before suspecting the proxy.
Timeouts once concurrency goes up
Lower CONCURRENT_REQUESTS_PER_DOMAIN and raise DOWNLOAD_TIMEOUT. RETRY_TIMES covers the odd slow residential peer without failing the crawl.
A wall of 403 or 429 responses
The site is limiting or blocking you. Turn on AUTOTHROTTLE_ENABLED, send a realistic USER_AGENT, and move from datacenter to residential if the blocks follow you from IP to IP.
Which line to pick
Crawls of sites that do not fight back run well on datacenter IPs. For marketplaces and search results that block server ranges, residential with Randomize IP gives requests fresh addresses. Residential is billed by the gigabyte, so keep the spider to the pages it parses.
Other setup guides
Not sure what a word means? The glossary explains it in plain English.
Stuck on a step?
Paste the command and the error into Discord, with your password taken out. People there have met most of these errors before.