Explainer29 September 20267 min read

Proxies for a student research project, done properly

Collecting web data for a thesis or coursework? When you need a proxy at all, how to stay polite and within the rules, and how to keep the cost small.

A lot of first-time proxy buyers are students with a dataset to build: prices across countries for an economics paper, public posts for a linguistics corpus, job adverts for a labour-market study. Academic research is on the allowed list of our acceptable use policy, and it is work we are glad to support. It also comes with obligations a hobby scraper does not have, and it is worth getting them right before your supervisor or an examiner asks.

Do you need a proxy at all?

Often not. Try without one first. Proxies earn their place in a research project for a few specific reasons:

  • Your university network gets blocked. Campus traffic leaves through a handful of addresses shared by thousands of people. A site that throttles one student's scraper can end up throttling the whole library.
  • Location is part of the method. If you are comparing what a site shows in different countries, you need an exit in each of them. Our datacenter and ISP lines let you choose the country, region and city; our residential line does not offer a country choice.
  • The site treats hosting networks differently from homes. Then you need to know that, and say so in your method section, because it affects what you collected.

If none of those apply, a polite scraper on your own connection is simpler, cheaper and easier to explain.

Ask before you collect

  • Your supervisor, and the ethics process. Many universities require ethics review for anything involving people's data, even when it is public. Ask early; a late approval can cost you a term.
  • Personal data. Names, usernames, photos and anything that could identify someone can be personal data under laws such as the GDPR, public or not. Collect only what the research question needs, and plan how you will store, pseudonymise and eventually delete it.
  • The site's terms. Read them. Some forbid automated access outright; some allow it for research; some offer an API or a data export that makes scraping unnecessary.
  • Profiles of individuals. Building a profile of a private person is banned under our policy whatever the purpose. Aggregate questions about many people are a different kind of study.

Robots.txt, rate and identity

robots.txt is the site's own statement of what automated clients may fetch and how often. It is not a law, but ignoring it in a research project is hard to defend. Python can read it for you:

import time
import urllib.robotparser

import requests

AGENT = "thesis-crawler ([email protected])"
PROXY = "http://USERNAME:PASSWORD@HOST:PORT"
PROXIES = {"http": PROXY, "https": PROXY}

robots = urllib.robotparser.RobotFileParser("https://www.example.org/robots.txt")
robots.read()
delay = robots.crawl_delay(AGENT) or 5


def polite_get(url, attempts=3):
    if not robots.can_fetch(AGENT, url):
        return None
    for _ in range(attempts):
        r = requests.get(url, headers={"User-Agent": AGENT}, proxies=PROXIES, timeout=30)
        time.sleep(delay)
        if r.status_code not in (429, 503):
            return r
        wait = r.headers.get("Retry-After", "")
        time.sleep(int(wait) if wait.isdigit() else delay * 10)
    return None

A few habits that go with it:

  • One request at a time unless you have a reason to go faster. A thesis rarely has a deadline measured in minutes.
  • Say who you are. A User-Agent with a contact address lets a site owner email you instead of blocking you. For research, being reachable is usually the ethical choice even when you use a proxy for location.
  • Back off when asked. A 429 or a 503 is the site telling you to slow down. The snippet above waits; do not work around it with more IPs.

Keep a record for the method section

Log every request with the URL, the time, the status code and which exit country it used. When you write up, you can state exactly what you collected, when, and from where, and someone else can repeat it. The log also shows your supervisor you stayed inside the limits you promised.

Keeping the cost small

  • Estimate the traffic before you order anything. How per-GB billing works has a formula and a script for it.
  • For a location study, a few static IPs, one per country, are often enough; ISP or datacenter explains which to pick.
  • Your balance does not expire, so what is left after the pilot is still there for the main collection. If the project ends early, the unused part can go back to the card or crypto wallet you paid from under our refund policy.

Prices are on the pricing page. If you are unsure whether your project fits, ask in Discord before you buy. We are glad to help you design a crawler you can defend in your viva.

Got a follow-up question?

Ask it in Discord. The answer helps whoever reads the thread next.

Join the Discorddiscord.gg/proxypanda
Start with $5Ask in Discord