Tutorial29 September 202610 min read

Web scraping with Node.js: a polite scraper with proxies

Web scraping with Node.js: a tested undici and cheerio scraper with proxy rotation, a per-host limit, Retry-After backoff, robots.txt checks and JSONL output.

Web scraping with Node.js needs three things: an HTTP client that can go through a proxy (undici's fetch with a ProxyAgent), an HTML parser (cheerio), and the manners that keep a scraper welcome. Those manners are a robots.txt check, a per-host concurrency limit, backoff that honours Retry-After, and stopping the moment a site shows a challenge page.

This post builds one file, scrape.mjs, that does all of that and writes one JSON object per line. We ran it end to end against a local test site through a local proxy that asks for a login, and the results are at the bottom. Nothing here is specific to one website: the selectors match a simple example page, and you swap in your own.

What you need

  • Node.js 22 or newer. Current undici releases refuse older versions.
  • Three packages, pinned to the versions we tested:
npm i [email protected] [email protected] [email protected]
  • HOST, PORT, USERNAME and PASSWORD from your service page in the dashboard. There is no shared gateway address, so every snippet uses those placeholders.
  • A site you are allowed to scrape. SITE_DOMAIN below stands for its domain. Read its terms before you start; are proxies legal? covers where the lines usually sit.

Paste the six blocks below, in order, into one file called scrape.mjs.

Step 1: the settings

import { createWriteStream } from 'node:fs';
import { setTimeout as sleep } from 'node:timers/promises';
import * as cheerio from 'cheerio';
import robotsParser from 'robots-parser';
import { fetch, ProxyAgent } from 'undici';

const SITE = 'SITE_DOMAIN';
const START_URL = `https://${SITE}/products?page=1`;
const PROXIES = ['http://USERNAME:PASSWORD@HOST:PORT'];
const NEW_TUNNEL_PER_REQUEST = true;
const USER_AGENT = 'my-price-notes/1.0 (+mailto:[email protected])';
const PER_HOST = 2;
const MIN_GAP_MS = 1000;
const MAX_ATTEMPTS = 4;
const MAX_WAIT_MS = 120_000;
const MAX_PAGES = 50;
const OUTPUT = 'products.jsonl';

const RETRY_STATUS = new Set([429, 500, 502, 503, 504]);
const CHALLENGE = /captcha|verify you are human|unusual traffic/i;

PER_HOST is how many requests may be in flight to one site at once, and MIN_GAP_MS is the smallest gap between two request starts to that site. MAX_WAIT_MS is the longest Retry-After the scraper will sit through before it gives up for the day. The user agent names your scraper and gives a way to reach you, which is both honest and what robots.txt rules are matched against.

Step 2: proxy rotation in Node.js

class Blocked extends Error {}
let stopReason = null;
function stop(message) {
  stopReason ??= message;
  return new Blocked(stopReason);
}

const agents = new Map();
let turn = 0;
function pickProxy() {
  const proxy = PROXIES[turn++ % PROXIES.length];
  if (NEW_TUNNEL_PER_REQUEST) return new ProxyAgent(proxy);
  if (!agents.has(proxy)) agents.set(proxy, new ProxyAgent(proxy));
  return agents.get(proxy);
}

Blocked marks the errors that should end the whole run, and stop records the first reason so every later request can see it.

How rotation works depends on the product. A residential service set to Randomize IP gives you one endpoint and a new exit IP for each new connection. The catch is that a ProxyAgent keeps its tunnel open and reuses it. In our test, one agent carried seven requests, three of them in parallel, through a single CONNECT, which on a rotating service means a single exit IP. So with NEW_TUNNEL_PER_REQUEST = true the code makes a fresh agent for every request and closes it afterwards. Our test proxy then logged one CONNECT per request.

ISP and datacenter services are a list of static IPs, and rotation is your job. List each one and turn fresh tunnels off:

const PROXIES = [
  'http://USERNAME:PASSWORD@HOST_1:PORT_1',
  'http://USERNAME:PASSWORD@HOST_2:PORT_2',
  'http://USERNAME:PASSWORD@HOST_3:PORT_3',
];
const NEW_TUNNEL_PER_REQUEST = false;

Each IP then gets one long-lived agent, and requests take turns across them. With three test proxies, each logged exactly one CONNECT for the whole run. If a password contains @, : or /, wrap it in encodeURIComponent first. The Node.js guide covers the ProxyAgent basics, the axios guide does the same for axios, and rotating vs sticky proxies explains when you want the address to stay put instead.

Step 3: a per-host concurrency limit

function hostLimiter(max, gapMs) {
  let active = 0;
  let nextStart = 0;
  const waiting = [];
  return async (task) => {
    if (active < max) active++;
    else await new Promise((resolve) => waiting.push(resolve));
    try {
      const start = Math.max(Date.now(), nextStart);
      nextStart = start + gapMs;
      await sleep(start - Date.now());
      return await task();
    } finally {
      const next = waiting.shift();
      if (next) next();
      else active--;
    }
  };
}

const limiters = new Map();
function limitFor(url, gapMs) {
  const host = new URL(url).host;
  if (!limiters.has(host)) limiters.set(host, hostLimiter(PER_HOST, gapMs));
  return limiters.get(host);
}

This is a small semaphore with a spacing rule, so there is no extra package to install. A request waits for a free slot, then waits until at least gapMs has passed since the previous start to the same host. When a slot frees up it is handed straight to the next waiting request, so the limit holds even under a burst.

With the gap set to zero, our test site never saw more than two requests at once. With the one-second gap, it saw one request start per second, which is the setting to keep. How many threads per proxy goes further into finding a safe number for a given site.

Step 4: retries that honour Retry-After

function retryAfterMs(value) {
  if (!value) return null;
  const seconds = Number(value);
  if (Number.isFinite(seconds)) return seconds * 1000;
  const date = Date.parse(value);
  return Number.isNaN(date) ? null : Math.max(0, date - Date.now());
}

const backoff = (attempt) => 1000 * 2 ** attempt * (0.5 + Math.random());

function rootCause(error) {
  while (error.cause) error = error.cause;
  return error.message;
}

async function get(url) {
  const dispatcher = pickProxy();
  try {
    const response = await fetch(url, {
      dispatcher,
      headers: { 'user-agent': USER_AGENT, accept: 'text/html,text/plain' },
      signal: AbortSignal.timeout(30_000),
    });
    return { status: response.status, headers: response.headers, body: await response.text() };
  } finally {
    if (NEW_TUNNEL_PER_REQUEST) await dispatcher.close();
  }
}

async function fetchPage(url) {
  for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
    if (stopReason) throw new Blocked(stopReason);
    let problem;
    let wait;
    try {
      const { status, headers, body } = await get(url);
      const isHtml = (headers.get('content-type') ?? '').includes('html');
      if (status === 403 || (isHtml && CHALLENGE.test(body))) throw stop(`challenge or 403 at ${url}`);
      if (status >= 200 && status < 300) return body;
      if (!RETRY_STATUS.has(status)) return null;
      problem = `HTTP ${status}`;
      wait = retryAfterMs(headers.get('retry-after')) ?? backoff(attempt);
    } catch (error) {
      if (error instanceof Blocked) throw error;
      problem = rootCause(error);
      wait = backoff(attempt);
    }
    if (attempt === MAX_ATTEMPTS) break;
    if (wait > MAX_WAIT_MS) throw stop(`asked to wait ${Math.round(wait / 1000)} s at ${url}`);
    console.warn(`${problem} on ${url}, retrying in ${Math.round(wait / 1000)} s`);
    await sleep(wait);
  }
  throw new Error(`gave up on ${url} after ${MAX_ATTEMPTS} attempts`);
}

What each outcome does:

  • 2xx: return the page.
  • 403, or an HTML page that looks like a captcha or "verify you are human": stop the whole run. More requests only dig the hole deeper; change something first.
  • 429, 500, 502, 503, 504, or a network error: wait and try again, up to MAX_ATTEMPTS. If the site sent Retry-After, the wait is exactly that, whether it came as seconds or as a date. Otherwise the wait doubles each time, with some randomness so parallel requests do not retry in lockstep.
  • A Retry-After longer than MAX_WAIT_MS: stop. A site asking you to come back in ten minutes means today's run is over.
  • Anything else, such as 404: skip that URL and carry on.

rootCause digs the real reason out of undici's fetch failed, such as Proxy response (407) !== 200 when HTTP Tunneling for a refused proxy login. Fixing 429 errors explains the rate limits behind most retries. On the cost side, our honesty page says: "Connection failures (timeouts, resets, and errors from our own gateway such as a 407 or a 502) are billed at zero and still shown, marked free." A 429 or a 5xx page from the site is traffic, so retrying one costs the bytes of that response as well as time. The stopping rules matter more than the retry count.

Step 5: check robots.txt

async function loadRobots(origin) {
  const url = `${origin}/robots.txt`;
  return robotsParser(url, (await fetchPage(url)) ?? '');
}

robots-parser reads the site's rules once at the start. isAllowed answers per URL, and getCrawlDelay feeds the gap used in step 6, so a site that asks for a slower crawl gets one. A missing robots.txt counts as "everything allowed", as the standard says. A robots.txt behind a 403 or a challenge stops the run before it starts.

Step 6: parse with cheerio and write JSONL

function parseProduct(html, url) {
  const $ = cheerio.load(html);
  return {
    url,
    title: $('h1').first().text().trim(),
    price: $('.price').first().text().trim() || null,
    scrapedAt: new Date().toISOString(),
  };
}

async function main() {
  const robots = await loadRobots(new URL(START_URL).origin);
  const gap = Math.max(MIN_GAP_MS, (robots.getCrawlDelay(USER_AGENT) ?? 0) * 1000);
  const allowed = (url) => robots.isAllowed(url, USER_AGENT) !== false;
  const out = createWriteStream(OUTPUT, { flags: 'a' });
  const seen = new Set();
  let saved = 0;
  let pageUrl = START_URL;
  let pages = 0;

  try {
    while (pageUrl && allowed(pageUrl) && pages++ < MAX_PAGES) {
      const html = await limitFor(pageUrl, gap)(() => fetchPage(pageUrl));
      if (!html) break;
      const $ = cheerio.load(html);
      const links = $('a.product-link')
        .map((_, a) => new URL($(a).attr('href'), pageUrl).href)
        .get()
        .filter((link) => allowed(link) && !seen.has(link));
      links.forEach((link) => seen.add(link));

      const results = await Promise.allSettled(
        links.map(async (link) => {
          const productHtml = await limitFor(link, gap)(() => fetchPage(link));
          if (!productHtml) return;
          out.write(JSON.stringify(parseProduct(productHtml, link)) + '\n');
          saved++;
        }),
      );
      for (const { status, reason } of results) {
        if (status === 'fulfilled') continue;
        if (reason instanceof Blocked) throw reason;
        console.warn(`Skipped: ${reason.message}`);
      }

      const next = $('a[rel="next"]').attr('href');
      pageUrl = next ? new URL(next, pageUrl).href : null;
    }
  } finally {
    out.end();
    console.log(`Saved ${saved} products to ${OUTPUT}`);
  }
}

try {
  await main();
} catch (error) {
  console.error(`Stopped: ${error.message}`);
  process.exitCode = 1;
} finally {
  await Promise.all([...agents.values()].map((agent) => agent.close()));
}

The crawl walks the listing pages by following rel="next", collects product links, drops any that robots.txt disallows or that were already seen, and fetches the products through the limiter. Each product becomes one line of JSON in products.jsonl. JSONL suits scraping: every line is complete on its own, a crash loses nothing already written, and a rerun appends. Promise.allSettled lets the products on a page finish even when one of them fails, and a Blocked result still stops the run.

The selectors (a.product-link, h1, .price) fit the example page. Open your target in a browser, inspect a product link and a price, and change those three strings.

Run it

node scrape.mjs

On our test site, which answered one listing page with a 429 and one product with a 503, the output was:

HTTP 429 on https://SITE_DOMAIN/products?page=2, retrying in 2 s
HTTP 503 on https://SITE_DOMAIN/products/5, retrying in 2 s
Saved 11 products to products.jsonl

And one line of products.jsonl:

{"url":"https://SITE_DOMAIN/products/1","title":"Item 1","price":"3.50","scrapedAt":"2026-09-30T10:28:22.198Z"}

How we tested it

We served a small HTTPS site on our own machine with three listing pages and twelve products, and a robots.txt that disallowed /private/ and asked for a one-second crawl delay. One listing page answered its first request with a 429 and Retry-After: 2; one product answered its first request with a 503; one product was a 404; one listing linked to a disallowed page. All traffic went through a local proxy that required a username and password.

  • The scraper saved 11 products, skipped the 404, and never requested the disallowed page.
  • It waited the full two seconds the 429 asked for, and retried the 503 after a backoff.
  • In residential mode the proxy saw one CONNECT per request; in static-list mode, one per IP.
  • With a challenge page served on one product, it stopped with Stopped: challenge or 403 at ..., exit code 1, and the four rows already saved stayed in the file.
  • A Retry-After: 600 stopped the run instead of sleeping for ten minutes.
  • A wrong proxy password gave Proxy response (407) !== 200 when HTTP Tunneling on robots.txt, then Stopped: gave up on ... after the retries.

When Node.js fetch is not enough

If the data you need is drawn by JavaScript after the page loads, fetch gets an empty shell and cheerio has nothing to parse. Check the network tab for a JSON endpoint the page calls, which is often easier to read than the HTML. If there is none, drive a browser: the Playwright and Puppeteer guides show both with a proxy. A browser moves far more data per page, which matters on a per-GB bill.

Quick answers

Which HTTP client should I use for web scraping with Node.js? undici's fetch with ProxyAgent, as here, or axios with https-proxy-agent. Node's built-in fetch has no option that takes a proxy URL.

Why does my rotating proxy show the same IP on every request? The client reused a kept-alive tunnel. Create a fresh ProxyAgent per request on a Randomize IP service.

Do I need p-limit? No. The small limiter in step 3 does the same job for one host, with a spacing rule on top.

Is cheerio enough for scraping? For pages whose HTML arrives complete, yes. For pages built by JavaScript, use a browser.

Next step

Point SITE_DOMAIN at a site you are allowed to scrape, set MAX_PAGES to 2, and run it once on a small top-up to see the pages, the timing and the bytes before you scale up. Rates for each product are on the pricing page. If something in the output puzzles you, paste it in Discord with the password removed.

Got a follow-up question?

Ask it in Discord. The answer helps whoever reads the thread next.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord