Setup guide
Crawl4AI proxy setup with ProxyConfig
Crawl4AI drives a Playwright browser and turns pages into markdown for language models. The proxy belongs to the browser settings, with the username and password in fields of their own. Routing a crawl through a proxy needs no LLM API key. Tested with Crawl4AI 0.9.4.
Before you start
- Python 3.10 or newer, which Crawl4AI requires.
- A recent Crawl4AI. Older examples pass a proxy string to BrowserConfig; that still runs in 0.9.4 but prints a deprecation warning.
- HOST, PORT, USERNAME and PASSWORD from the service page.
Your connection details
Sign in to the dashboard and open the service you bought. Its page shows the host, port, username and password for that service. There is no single ProxyPanda address to remember, so copy them from there each time. The code in these guides uses the placeholders below; replace each one with your own value.
- HOST
- The proxy address for this service, exactly as the service page shows it.
- PORT
- The port to connect to. Copy it with the host, since it can differ from one service to the next.
- USERNAME
- Your proxy login. It is separate from the email you use for the dashboard.
- PASSWORD
- Copy it in full. On residential, the options you choose in the dashboard are added to the end of the password, so a password typed from memory loses them.
You can skip the username and password by adding the IP address you connect from to the service’s allowlist. Then only HOST and PORT go into your code.
The examples use HTTP, which reaches https sites through an encrypted tunnel. SOCKS5 works too: residential ports accept both, and ISP and datacenter proxies switch protocol in the dashboard.
Steps
Install Crawl4AI in a virtual environment
crawl4ai-setup downloads the Chromium build that Playwright drives and prepares Crawl4AI’s local database. Keeping it all in .venv stops one project’s browser and packages from colliding with another’s.
terminal python3 -m venv .venv source .venv/bin/activate pip install -U crawl4ai crawl4ai-setupGive BrowserConfig a ProxyConfig
server carries the scheme, host and port and nothing more. username and password sit in separate fields, and the browser answers the proxy’s login challenge with them. Every page this crawler opens goes through the proxy.
crawl_ip.py import asyncio from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, ProxyConfig proxy = ProxyConfig(server="http://HOST:PORT", username="USERNAME", password="PASSWORD") browser_config = BrowserConfig(headless=True, proxy_config=proxy) async def main(): async with AsyncWebCrawler(config=browser_config) as crawler: result = await crawler.arun("https://api.ipify.org?format=json", config=CrawlerRunConfig()) print(result.success, result.status_code) print(result.markdown) asyncio.run(main())Read what came back
The script prints True 200, then the JSON from api.ipify.org wrapped in a markdown code block. The ip value is the proxy’s exit address.
Rotate across a list with RoundRobinProxyStrategy
ProxyConfig.from_string reads a line in HOST:PORT:USERNAME:PASSWORD order, the order the dashboard exports. The strategy gives each arun call the next entry, and the log prints a Switch proxy line every time. page_timeout is in milliseconds.
rotate.py import asyncio from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, ProxyConfig, RoundRobinProxyStrategy with open("proxies.txt") as f: proxies = [ProxyConfig.from_string(line) for line in f if line.strip()] run_config = CrawlerRunConfig( proxy_rotation_strategy=RoundRobinProxyStrategy(proxies), check_robots_txt=True, page_timeout=60000, ) async def main(): urls = ["https://api.ipify.org?format=json"] * 6 async with AsyncWebCrawler() as crawler: for url in urls: result = await crawler.arun(url, config=run_config) print(result.success, result.markdown if result.success else result.error_message) asyncio.run(main())Keep the crawl polite
An AI crawler is still a crawler: robots.txt and a site’s terms apply to it as to any scraper, and a steady pace is kinder to the site than a burst. check_robots_txt=True skips disallowed URLs. In 0.9.4 that robots.txt request goes out directly from your machine, not through the proxy, with a two-second limit, and a file it cannot fetch counts as allowing everything.
Rotating and sticky IPs
Crawl4AI keeps a browser context per proxy entry and reuses it. In testing, two crawls through the same entry shared one tunnel, and a shared tunnel leaves from one address. On residential with Randomize IP, a new exit arrives with a new connection, which is not the same as with every page.
For a sign-in followed by a few pages, use Sticky IP on residential or a static ISP IP. Setting proxy_session_id on CrawlerRunConfig makes RoundRobinProxyStrategy keep one entry from the list for every crawl that shares the id.
Wide crawls of public pages suit Randomize IP, since the requests spread across many residential addresses.
Common errors and fixes
Page.goto: Timeout 60000ms exceeded
A wrong username or password shows up this way: Chromium keeps retrying the proxy login until the page times out, and no 407 is reported. Check the credentials before raising page_timeout.
net::ERR_PROXY_CONNECTION_FAILED
The browser could not reach HOST:PORT. Compare both with the service page and confirm the service is active.
Blocked by anti-bot protection: Structural: minimal_text on small page
Crawl4AI’s own block detector marked a page with almost no text as blocked, even though the site answered 200. The plain-text reply from api.ipify.org trips it; the ?format=json address used above passes.
ValueError: Invalid proxy string format
from_string splits a line on colons and expects four parts. A password containing a colon breaks that; build that entry with ProxyConfig(server=..., username=..., password=...) instead.
UserWarning: The 'proxy' parameter is deprecated
The code passes BrowserConfig(proxy=...) from an older tutorial. Move the value into proxy_config as in step 2.
Which line to pick
Crawls that feed a model often read many pages once each. Datacenter IPs keep that cheap on sites that accept server traffic. Use residential for sites that block data-centre ranges, and keep pages light while you do: residential is billed by the gigabyte, and text_mode=True on BrowserConfig tells the browser to skip images.
Other setup guides
Not sure what a word means? The glossary explains it in plain English.
Stuck on a step?
Paste the command and the error into Discord, with your password taken out. People there have met most of these errors before.