Tutorial29 September 20269 min read

How to scrape Amazon prices with Python

Scrape Amazon prices from public product pages with Python, requests and BeautifulSoup. Parse title and price by ASIN, spot a robot check, pace it, save to CSV.

To scrape Amazon prices, fetch each public product page by its ASIN at a slow, steady pace, parse the title and price with BeautifulSoup using several fallback selectors, stop the run the moment a robot-check page appears, and append each result to a CSV file. A static IP in the marketplace's country keeps the numbers comparable from one day to the next.

This tutorial builds that tracker in about a hundred lines of Python. It covers what is reasonable to collect, which proxy line to start on, the parser, pacing, how to tell a robot check from a product page, and how much traffic the job will use.

Is it OK to scrape Amazon prices?

Be clear about this before you write any code. Amazon's Conditions of Use do not allow robots, data mining or similar data extraction tools on the site. A price tracker is one of those, even when every page it reads is public. What follows is a way to collect public prices politely. It does not make the activity something Amazon has agreed to, and whether terms like these bind you depends on where you live and how you use the data. Are proxies legal? covers the general picture; for anything commercial, ask someone qualified.

If you want to go ahead, keep to these limits:

  • Public product pages only. No login, no account, no cart, nothing behind a sign-in. Never scrape while signed in to your own Amazon account.
  • A low rate. A few hundred products once or twice a day is a tracker. Thousands of pages an hour is a load on someone else's site.
  • Read the marketplace's robots.txt and keep away from the paths it disallows.
  • Consider the official route first. If you are a seller or an affiliate, Amazon offers APIs that return price data without any of this.

Price tracking of prices any visitor can see is fine on our side, and the price monitoring use case describes it. Creating accounts to reach member-only prices is not.

Which proxy line fits an Amazon price scraper?

Amazon's prices, delivery options and sometimes the offer it shows depend on two things: which marketplace domain you visit, and the delivery location it guesses from your IP address. A tracker should ask from the same country every time, or yesterday's price and today's price are not the same measurement.

That points at the static lines:

  • Start with datacenter. You pick the country, region and city when you order, and it is the cheapest per IP. Test it on a handful of ASINs for a few days.
  • Move to ISP if datacenter keeps hitting robot checks even at a slow pace. ISP IPs are registered to consumer internet providers, and shops wary of server traffic tend to treat them more like a shopper at home. You still choose the location. ISP or datacenter proxies? compares the two.
  • Residential is the poor fit here. It rotates through home connections and has no country, region or city choice, so two checks of the same product can come from two different countries.

The locations pages list where static IPs are available. Residential or datacenter? has a five-minute test for deciding on your own target.

The Amazon price scraper in Python

Install the two libraries:

pip install requests beautifulsoup4

Then save this as amazon_prices.py. Put your ASINs in ASINS, set MARKETPLACE to the marketplace's domain, and fill in the proxy values from your service page.

import csv
import os
import random
import re
import time
from datetime import datetime, timezone

import requests
from bs4 import BeautifulSoup

PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
MARKETPLACE = "MARKETPLACE_DOMAIN"
ASINS = ["ASIN_1", "ASIN_2", "ASIN_3"]
OUTPUT = "prices.csv"

HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
}

TITLE_SELECTORS = ["#productTitle", "#title"]
PRICE_SELECTORS = [
    "#corePriceDisplay_desktop_feature_div .a-price .a-offscreen",
    "#corePrice_feature_div .a-price .a-offscreen",
    "#apex_desktop .a-price .a-offscreen",
    "#price_inside_buybox",
    "#priceblock_ourprice",
    "#priceblock_dealprice",
]
ROBOT_MARKERS = ("validatecaptcha", "type the characters you see", "robot check")


class Blocked(Exception):
    pass


def first_text(soup, selectors):
    for selector in selectors:
        node = soup.select_one(selector)
        if node and node.get_text(strip=True):
            return node.get_text(strip=True)
    return None


def to_number(text):
    digits = re.sub(r"[^\d.,]", "", text or "")
    if not digits:
        return None
    last = max(digits.rfind("."), digits.rfind(","))
    if last != -1 and len(digits) - last - 1 == 2:
        whole, fraction = digits[:last], digits[last + 1:]
    else:
        whole, fraction = digits, "0"
    return float(re.sub(r"[.,]", "", whole) + "." + fraction)


def parse_product(html):
    if any(marker in html.lower() for marker in ROBOT_MARKERS):
        raise Blocked("robot check page")
    soup = BeautifulSoup(html, "html.parser")
    price_text = first_text(soup, PRICE_SELECTORS)
    return {
        "title": first_text(soup, TITLE_SELECTORS),
        "price_text": price_text,
        "price": to_number(price_text),
    }


def fetch(asin):
    url = f"https://{MARKETPLACE}/dp/{asin}"
    response = requests.get(url, headers=HEADERS, proxies=PROXIES, timeout=(5, 30))
    if response.status_code in (429, 503):
        raise Blocked(f"HTTP {response.status_code}")
    if response.status_code == 404:
        return None
    response.raise_for_status()
    item = parse_product(response.text)
    if item["title"] is None:
        with open(f"unparsed-{asin}.html", "w", encoding="utf-8") as f:
            f.write(response.text)
    return item


def main():
    new_file = not os.path.exists(OUTPUT)
    with open(OUTPUT, "a", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        if new_file:
            writer.writerow(["checked_at", "marketplace", "asin", "title", "price_text", "price"])
        for asin in ASINS:
            try:
                item = fetch(asin)
            except Blocked as reason:
                print(f"Stopping at {asin}: {reason}")
                break
            except requests.RequestException as error:
                print(f"{asin}: {error}")
                item = None
            if item:
                writer.writerow([
                    datetime.now(timezone.utc).isoformat(timespec="seconds"),
                    MARKETPLACE, asin, item["title"], item["price_text"], item["price"],
                ])
                print(asin, item["price_text"], item["title"])
            time.sleep(random.uniform(8, 15))


if __name__ == "__main__":
    main()

What each part does

  • The URL is built from the marketplace and the ASIN, the ten-character product code in every Amazon product link after /dp/. /dp/ASIN is the shortest stable form of a product page.
  • PRICE_SELECTORS is a list, tried in order. Amazon has moved the main price between several blocks over the years, and different product types still use different ones. The first selector that finds text wins. The list deliberately avoids a bare .a-price catch-all, because that also matches prices in "customers also bought" carousels and would record the wrong product.
  • to_number keeps the raw text and a parsed number. Marketplaces write prices differently: 1,299.00 in one, 1.299,00 in another, and whole numbers with no decimals in a third. The function treats the last separator as the decimal point only when exactly two digits follow it. The raw price_text goes into the CSV too, so you can always check the parse.
  • A missing title means an unexpected page. The HTML is saved to unparsed-ASIN.html so you can see what came back, which is how you notice a layout change before a week of empty rows.
  • A missing price is not an error. Unavailable products, and products sold only through other sellers' offers, often have no main price. The row is written with an empty price, which is itself useful history.

Selectors will change

Amazon's markup changes often and without notice, and nothing in this list is guaranteed next month. When rows start coming back empty, open a product page in your browser, right-click the price, choose Inspect, and look for the nearest element with a stable id. Add its selector to the front of PRICE_SELECTORS.

Test the change on the saved file before spending any traffic:

from amazon_prices import parse_product

with open("unparsed-ASIN_1.html", encoding="utf-8") as f:
    print(parse_product(f.read()))

Headers and pacing

  • User-Agent: copy the string from your own browser. A default python-requests agent announces a script on the first request.
  • Accept-Language: match the marketplace. en-US for a US store, de-DE for a German one, and so on.
  • Pace: the script waits 8 to 15 seconds between products, with some randomness so the rhythm does not look mechanical. Do not run several copies in parallel on one IP; if you need more throughput, add IPs and give each its own share of the ASIN list.
  • Schedule: most prices do not change by the minute. Once or twice a day from cron or Task Scheduler is plenty for most tracking.

When the robot check appears, stop

Amazon answers suspected automation with a page asking you to type the characters from an image, sometimes with status 200 and sometimes with 503. parse_product looks for the markers of that page, and the loop stops at the first one instead of carrying on.

Stopping is the point. A robot check means the site has doubts about this IP or this traffic, and every further request adds to them. Hammering on through retries turns a short cool-down into a lasting block and, on a per-GB line, pays for each challenge page. Wait several hours, lower the pace, and try again. If checks keep appearing at a gentle rate, that is the signal to move from datacenter to ISP. Why your scraper started getting blocked lists the other usual causes, and retrying without burning bandwidth covers backoff for the transient errors.

How much traffic will an Amazon price tracker use?

Amazon product pages are heavy HTML, even without images, which this script never downloads. Measure your own before you plan: how per-GB billing works has a script that weighs a sample of pages through the proxy.

A worked example with made-up round numbers: if a product page weighs 150 KB on the wire and you track 200 ASINs twice a day for 30 days, that is 12,000 pages, or about 1.8 GB. Datacenter and ISP IPs are rented per IP for a term, and the pricing page shows the traffic each IP plan includes, so check that the job fits.

Reading the CSV

Each run appends rows, so prices.csv becomes a price history you can open in any spreadsheet. To print only the products whose price moved since the previous check:

import csv
from collections import defaultdict

history = defaultdict(list)
with open("prices.csv", newline="", encoding="utf-8") as f:
    for row in csv.DictReader(f):
        if row["price"]:
            history[row["asin"]].append(float(row["price"]))

for asin, prices in history.items():
    if len(prices) > 1 and prices[-1] != prices[-2]:
        print(asin, prices[-2], "->", prices[-1])

Quick answers

Why does my scraper see a different price from my browser? Your browser has your delivery address, your sign-in and your cookies. The scraper sees what a signed-out visitor in the proxy's country sees. Compare against a private browser window with no sign-in.

Can I scrape Amazon prices without a proxy? For a few products, occasionally, yes. Repeated checks from your home or server address will draw robot checks sooner, and you only ever see your own country's view.

Why is the price empty for some ASINs? The product may be unavailable, sold only through other sellers' offers, or on a layout your selectors do not cover. Open it in a browser to see which.

Can I run it every few minutes? You can, and it is a good way to get the IP flagged. Prices move slowly; check at the pace they change.

Start small

One datacenter IP in the marketplace's country is enough to run this against a few dozen products for a week and see whether the pages come back clean. Rates for each term are on the pricing page, and if a target gives you trouble, describe it in Discord before you spend more.

Got a follow-up question?

Ask it in Discord. The answer helps whoever reads the thread next.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord