Before you start

  • Python 3.10 or newer, which Crawl4AI requires.
  • A recent Crawl4AI. Older examples pass a proxy string to BrowserConfig; that still runs in 0.9.4 but prints a deprecation warning.
  • HOST, PORT, USERNAME and PASSWORD from the service page.

Your connection details

Sign in to the dashboard and open the service you bought. Its page shows the host, port, username and password for that service. There is no single ProxyPanda address to remember, so copy them from there each time. The code in these guides uses the placeholders below; replace each one with your own value.

HOST
The proxy address for this service, exactly as the service page shows it.
PORT
The port to connect to. Copy it with the host, since it can differ from one service to the next.
USERNAME
Your proxy login. It is separate from the email you use for the dashboard.
PASSWORD
Copy it in full. On residential, the options you choose in the dashboard are added to the end of the password, so a password typed from memory loses them.

You can skip the username and password by adding the IP address you connect from to the service’s allowlist. Then only HOST and PORT go into your code.

The examples use HTTP, which reaches https sites through an encrypted tunnel. SOCKS5 works too: residential ports accept both, and ISP and datacenter proxies switch protocol in the dashboard.

Steps

  1. Install Crawl4AI in a virtual environment

    crawl4ai-setup downloads the Chromium build that Playwright drives and prepares Crawl4AI’s local database. Keeping it all in .venv stops one project’s browser and packages from colliding with another’s.

    terminal
    python3 -m venv .venv
    source .venv/bin/activate
    pip install -U crawl4ai
    crawl4ai-setup
  2. Give BrowserConfig a ProxyConfig

    server carries the scheme, host and port and nothing more. username and password sit in separate fields, and the browser answers the proxy’s login challenge with them. Every page this crawler opens goes through the proxy.

    crawl_ip.py
    import asyncio
    
    from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, ProxyConfig
    
    proxy = ProxyConfig(server="http://HOST:PORT", username="USERNAME", password="PASSWORD")
    browser_config = BrowserConfig(headless=True, proxy_config=proxy)
    
    
    async def main():
        async with AsyncWebCrawler(config=browser_config) as crawler:
            result = await crawler.arun("https://api.ipify.org?format=json", config=CrawlerRunConfig())
            print(result.success, result.status_code)
            print(result.markdown)
    
    
    asyncio.run(main())
  3. Read what came back

    The script prints True 200, then the JSON from api.ipify.org wrapped in a markdown code block. The ip value is the proxy’s exit address.

  4. Rotate across a list with RoundRobinProxyStrategy

    ProxyConfig.from_string reads a line in HOST:PORT:USERNAME:PASSWORD order, the order the dashboard exports. The strategy gives each arun call the next entry, and the log prints a Switch proxy line every time. page_timeout is in milliseconds.

    rotate.py
    import asyncio
    
    from crawl4ai import AsyncWebCrawler, CrawlerRunConfig, ProxyConfig, RoundRobinProxyStrategy
    
    with open("proxies.txt") as f:
        proxies = [ProxyConfig.from_string(line) for line in f if line.strip()]
    
    run_config = CrawlerRunConfig(
        proxy_rotation_strategy=RoundRobinProxyStrategy(proxies),
        check_robots_txt=True,
        page_timeout=60000,
    )
    
    
    async def main():
        urls = ["https://api.ipify.org?format=json"] * 6
        async with AsyncWebCrawler() as crawler:
            for url in urls:
                result = await crawler.arun(url, config=run_config)
                print(result.success, result.markdown if result.success else result.error_message)
    
    
    asyncio.run(main())
  5. Keep the crawl polite

    An AI crawler is still a crawler: robots.txt and a site’s terms apply to it as to any scraper, and a steady pace is kinder to the site than a burst. check_robots_txt=True skips disallowed URLs. In 0.9.4 that robots.txt request goes out directly from your machine, not through the proxy, with a two-second limit, and a file it cannot fetch counts as allowing everything.

Rotating and sticky IPs

Crawl4AI keeps a browser context per proxy entry and reuses it. In testing, two crawls through the same entry shared one tunnel, and a shared tunnel leaves from one address. On residential with Randomize IP, a new exit arrives with a new connection, which is not the same as with every page.

For a sign-in followed by a few pages, use Sticky IP on residential or a static ISP IP. Setting proxy_session_id on CrawlerRunConfig makes RoundRobinProxyStrategy keep one entry from the list for every crawl that shares the id.

Wide crawls of public pages suit Randomize IP, since the requests spread across many residential addresses.

Common errors and fixes

Page.goto: Timeout 60000ms exceeded

A wrong username or password shows up this way: Chromium keeps retrying the proxy login until the page times out, and no 407 is reported. Check the credentials before raising page_timeout.

net::ERR_PROXY_CONNECTION_FAILED

The browser could not reach HOST:PORT. Compare both with the service page and confirm the service is active.

Blocked by anti-bot protection: Structural: minimal_text on small page

Crawl4AI’s own block detector marked a page with almost no text as blocked, even though the site answered 200. The plain-text reply from api.ipify.org trips it; the ?format=json address used above passes.

ValueError: Invalid proxy string format

from_string splits a line on colons and expects four parts. A password containing a colon breaks that; build that entry with ProxyConfig(server=..., username=..., password=...) instead.

UserWarning: The 'proxy' parameter is deprecated

The code passes BrowserConfig(proxy=...) from an older tutorial. Move the value into proxy_config as in step 2.

Which line to pick

Crawls that feed a model often read many pages once each. Datacenter IPs keep that cheap on sites that accept server traffic. Use residential for sites that block data-centre ranges, and keep pages light while you do: residential is billed by the gigabyte, and text_mode=True on BrowserConfig tells the browser to skip images.

Stuck on a step?

Paste the command and the error into Discord, with your password taken out. People there have met most of these errors before.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord