How to scrape eBay listings with Python
Scrape eBay listings from public search pages with Python. Parse title, price, condition and shipping, track sold prices, stop at challenge pages, save to CSV.
To scrape eBay listings, request public search result pages for your keywords at a slow pace, parse each listing's title, price, condition and shipping with fallback selectors, stop the run as soon as a challenge page appears, and append the rows to a CSV. A static IP in the marketplace's country keeps prices and shipping comparable between runs.
This tutorial builds that scraper with requests and BeautifulSoup, adds a sold-prices mode, and covers which proxy line to use and how much traffic it needs. First, one thing needs saying plainly.
Is it OK to scrape eBay?
eBay's User Agreement says you may not use any robot, spider, scraper, data mining or data extraction tool, or other automated means to access its services for any purpose without eBay's prior express permission. That covers a script like this one, even though it only reads public search pages. Following the steps below makes your scraping polite; it does not make it permitted. Whether terms like these bind you depends on where you live and what you do with the data, and are proxies legal? covers the general picture. For anything commercial, ask someone qualified.
The official route comes first. eBay's developer program offers the Browse API, which searches active listings and returns titles, prices, conditions and shipping as JSON, with no HTML to parse and nothing to break when the page design changes. As far as we can tell, sold-listing data is only available through a separate API with restricted access. If the API covers your use, use it.
If you still want to read the public pages for a small personal tracker, keep to these limits:
- Public search pages only. Never sign in, and never scrape from your own eBay buying or selling account.
- A low rate. A few dozen searches once a day is a tracker. The script waits 10 to 20 seconds between pages.
- Read the marketplace's robots.txt and keep away from the paths it disallows.
- No bidding, buying or messaging from code. eBay's agreement singles out automated ordering, and it is not something we support either.
Tracking prices anyone can see is on the allowed side of our price monitoring use case.
Which proxy line fits an eBay price tracker?
What eBay shows depends on which marketplace domain you open and where it thinks you are: shipping costs and delivery estimates are worked out for a location, and a listing that does not ship there may be shown differently or left out. A tracker should always ask from the same country.
- Start with datacenter. You choose the country, region and city when you order, and it is the cheapest per IP. Run it for a few days on a handful of searches.
- Move to ISP if challenge pages keep appearing even at a slow pace. ISP IPs are registered to consumer internet providers, and marketplaces tend to treat them like shoppers at home. ISP or datacenter proxies? compares the two.
- Residential is the poor fit. It has no country choice, so two runs of the same search can come from two countries and show different shipping.
Static IPs are available in the countries on the locations pages.
Scrape eBay listings with Python: the script
Install the libraries:
pip install requests beautifulsoup4
Save this as ebay_listings.py. Set EBAY to the marketplace domain, put your search terms in QUERIES, and fill in the proxy values from your service page.
import csv
import os
import random
import re
import time
from datetime import datetime, timezone
from urllib.parse import urlencode
import requests
from bs4 import BeautifulSoup
PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}
EBAY = "EBAY_DOMAIN"
QUERIES = ["QUERY_1", "QUERY_2"]
SOLD_ONLY = False
MAX_PAGES = 2
OUTPUT = "ebay_listings.csv"
HEADERS = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.9",
}
ITEM_SELECTORS = ["li.s-card", "li.s-item"]
TITLE_SELECTORS = [".s-card__title", ".s-item__title"]
PRICE_SELECTORS = [".s-card__price", ".s-item__price"]
CONDITION_SELECTORS = [".s-card__subtitle", ".s-item__subtitle .SECONDARY_INFO"]
SHIPPING_SELECTORS = [".s-item__shipping", ".s-item__logisticsCost"]
CAPTION_SELECTORS = [".s-card__caption", ".s-item__caption--signal", ".s-item__caption"]
NOISE_SELECTORS = ".clipped, .LIGHT_HIGHLIGHT"
SHIPPING_WORDS = re.compile(r"\b(shipping|delivery|postage)\b", re.I)
ITEM_ID = re.compile(r"/itm/(?:[^/?#]+/)?(\d{9,})")
CHALLENGE_MARKERS = ("pardon our interruption", "splashui/captcha", "checking your browser")
class Blocked(Exception):
pass
def search_url(query, page):
params = {"_nkw": query, "_pgn": page}
if SOLD_ONLY:
params.update(LH_Sold=1, LH_Complete=1)
return f"https://{EBAY}/sch/i.html?{urlencode(params)}"
def first_text(node, selectors):
for selector in selectors:
found = node.select_one(selector)
if found and found.get_text(strip=True):
return " ".join(found.get_text(" ", strip=True).split())
return None
def shipping_text(item, title):
text = first_text(item, SHIPPING_SELECTORS)
if text:
return text
for line in item.stripped_strings:
if SHIPPING_WORDS.search(line) and line not in (title or ""):
return line
return None
def low_price(text):
match = re.search(r"\d[\d.,]*", text or "")
if not match:
return None
digits = match.group().rstrip(".,")
last = max(digits.rfind("."), digits.rfind(","))
if last != -1 and len(digits) - last - 1 == 2:
whole, fraction = digits[:last], digits[last + 1:]
else:
whole, fraction = digits, "0"
return float(re.sub(r"[.,]", "", whole) + "." + fraction)
def parse_results(html):
if any(marker in html.lower() for marker in CHALLENGE_MARKERS):
raise Blocked("challenge page")
soup = BeautifulSoup(html, "html.parser")
for noise in soup.select(NOISE_SELECTORS):
noise.decompose()
items = []
for selector in ITEM_SELECTORS:
items = soup.select(selector)
if items:
break
rows = []
for item in items:
link = item.select_one("a[href*='/itm/']")
found = ITEM_ID.search(link["href"]) if link else None
if not found:
continue
title = first_text(item, TITLE_SELECTORS)
price_text = first_text(item, PRICE_SELECTORS)
rows.append({
"item_id": found.group(1),
"title": title,
"price_text": price_text,
"price": low_price(price_text),
"condition": first_text(item, CONDITION_SELECTORS),
"shipping": shipping_text(item, title),
"caption": first_text(item, CAPTION_SELECTORS),
})
return rows
def fetch_page(query, page):
response = requests.get(search_url(query, page), headers=HEADERS, proxies=PROXIES, timeout=(5, 30))
if response.status_code in (429, 503) or "/splashui/" in response.url:
raise Blocked(f"HTTP {response.status_code} at {response.url}")
response.raise_for_status()
rows = parse_results(response.text)
if not rows and page == 1:
name = re.sub(r"\W+", "-", query)
with open(f"unparsed-{name}.html", "w", encoding="utf-8") as f:
f.write(response.text)
return rows
def main():
new_file = not os.path.exists(OUTPUT)
seen = set()
with open(OUTPUT, "a", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
if new_file:
writer.writerow(["checked_at", "query", "item_id", "title", "price_text",
"price", "condition", "shipping", "caption"])
for query in QUERIES:
for page in range(1, MAX_PAGES + 1):
try:
rows = fetch_page(query, page)
except Blocked as reason:
print(f"Stopping at '{query}' page {page}: {reason}")
return
except requests.RequestException as error:
print(f"'{query}' page {page}: {error}")
break
checked_at = datetime.now(timezone.utc).isoformat(timespec="seconds")
for row in rows:
if row["item_id"] in seen:
continue
seen.add(row["item_id"])
writer.writerow([checked_at, query, row["item_id"], row["title"], row["price_text"],
row["price"], row["condition"], row["shipping"], row["caption"]])
print(f"'{query}' page {page}: {len(rows)} listings")
time.sleep(random.uniform(10, 20))
if not rows:
break
if __name__ == "__main__":
main()
What each part does
search_urlbuilds a normal search page address:_nkwis the search text and_pgnthe page number. WithSOLD_ONLY = Trueit addsLH_Sold=1&LH_Complete=1, the same filters as the "Sold items" option in the search filters.- Two layouts, one parser. eBay's search results have used an older
s-itemmarkup and, more recently, ans-cardone, and you may meet either.ITEM_SELECTORStries the newer one first, and every field has a list of selectors tried in order. - The item number is the key. Each listing links to
/itm/followed by a long item number. Promotional tiles and placeholders without a real item number are skipped, and a listing that shows up on two pages is written once. - Noise is removed before reading text. Screen-reader labels such as "Opens in a new window or tab" and "New listing" badges sit inside the title element;
NOISE_SELECTORSdeletes them first. - Shipping has a text fallback. When no known shipping element is present, the parser takes the first line in the listing that mentions shipping, delivery or postage, skipping the title so a "delivery bag" in a product name is not mistaken for a shipping cost.
- Prices keep their raw text.
price_textgoes into the CSV as shown.priceis the first number in it, so a range such as "10.00 to 15.00" records its low end, and both1,299.00and1.299,00read as 1299. - An empty first page is saved. No results on page 1 means a search with no matches, or a layout the selectors do not know. The HTML goes to
unparsed-QUERY.htmlso you can tell which.
We tested it through a password-protected local proxy on fixture pages in both layouts, with placeholders, a price range, sold captions and European number formats. Real pages will differ in the details, which is why the next section exists.
When the selectors stop matching
eBay's markup changes without notice, and the class names above are the ones in use when this was written. When a run starts returning empty fields, open a search page in your browser, right-click a price, choose Inspect, and add the new class to the front of the matching list. Test on a saved page before spending any traffic:
from ebay_listings import parse_results
with open("unparsed-QUERY_1.html", encoding="utf-8") as f:
for row in parse_results(f.read()):
print(row)
Detect challenge pages, and stop
When eBay suspects automation, it can answer with a page titled "Pardon Our Interruption", a redirect to a captcha under /splashui/, a 429 or a 503. parse_results looks for the markers of the first two, and fetch_page checks the status and the final address. Any of them raises Blocked, and the run ends.
Stopping is deliberate: a challenge means the site has doubts about this IP or this traffic, and every further request adds to them. Wait several hours and slow down. If challenges keep coming at a gentle pace from a datacenter IP, try ISP. Why your scraper started getting blocked goes through the other usual causes, and fixing 429 Too Many Requests covers rate limits.
Tracking eBay sold prices
Asking prices tell you what sellers hope for. Sold prices tell you what buyers paid, which is what most people want from an eBay price tracker. Set SOLD_ONLY = True, and each row gets a caption such as "Sold Sep 12, 2026". A few lines turn the CSV into a summary per search:
import csv
import statistics
from collections import defaultdict
sold = defaultdict(list)
with open("ebay_listings.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
if row["price"] and (row["caption"] or "").startswith("Sold"):
sold[row["query"]].append(float(row["price"]))
for query, prices in sold.items():
print(f"{query}: {len(prices)} sold, median {statistics.median(prices):.2f}, "
f"lowest {min(prices):.2f}, highest {max(prices):.2f}")
Use the median rather than the average: one mislabelled bundle or a broken item sold for parts can drag an average a long way. Keep one marketplace per CSV, since a single file mixing domains mixes currencies too.
How much traffic does an eBay scraper use?
Search pages are heavier than they look, even without images, which this script never downloads. Measure a few of your own through the proxy first; how per-GB billing works has a script for it.
A worked example with made-up round numbers: if a results page weighs 400 KB and you check 20 searches, two pages each, once a day for 30 days, that is 1,200 pages, or about 480 MB. Datacenter and ISP IPs are rented per IP for a term, and the pricing page shows the traffic each plan includes.
Quick answers
Why do my scraped prices differ from what I see in the browser? Your browser has your location, sign-in and cookies. The script sees what a signed-out visitor in the proxy's country sees. Compare against a private window.
Can I scrape eBay without a proxy? For a couple of searches, occasionally. Repeated runs from your own address meet challenges sooner, and you only ever see your own country's shipping.
Should I scrape item pages too? Only if the search page lacks a field you need. One item page per listing multiplies the requests many times over, and the API returns item details directly.
Start small
One datacenter IP in the marketplace's country is enough to run this against a few searches for a week and see whether pages come back clean. Rates are on the pricing page, and if a marketplace gives you trouble, describe it in Discord before you spend more.