A crawler starts from one page, collects the links on it, visits those, and keeps going until it has seen what it needs. With asyncio and httpx, a handful of workers can fetch many pages at once while the proxy gives every request its own IP. This guide builds a complete crawler step by step; the full script is at the end.
You need Python 3.10+ and httpx:
pip install httpxYour proxy username and password are in the generator. The examples use the Residential gateway geo.crawlproxies.com:8080 and crawl books.toscrape.com, a sandbox site built for scraping practice.
The proxy client
httpx takes the proxy as one URL, credentials included:
import httpx
PROXY = "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
limits = httpx.Limits(max_connections=8, max_keepalive_connections=0)
client = httpx.AsyncClient(proxy=PROXY, limits=limits, timeout=30, follow_redirects=True)max_keepalive_connections=0 is what makes the IPs rotate. With a rotating username, every new connection gets a new IP, and httpx normally keeps connections open for the next request. With keep-alive off, each request opens its own connection and leaves from its own address. max_connections caps how many requests run at once.
Using an older httpx? Versions before 0.26 call the argumentproxies=instead ofproxy=.
Finding links
The standard library's HTMLParser is enough to pull hrefs out of a page:
from html.parser import HTMLParser
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag == "a":
href = dict(attrs).get("href")
if href:
self.links.append(href)Links are relative more often than not, so join each one with the page's URL and drop the #fragment part, or the same page gets crawled twice:
from urllib.parse import urldefrag, urljoin
link = urldefrag(urljoin(page_url, href)).urlRetries with backoff
Some requests fail: a slow exit, a 429, a 503. Retrying after a short, growing pause fixes most of them, and because keep-alive is off, every retry also comes from a new IP:
import asyncio
import random
RETRY = {429, 500, 502, 503, 504}
async def fetch(client, url):
for attempt in range(4):
try:
r = await client.get(url)
if r.status_code not in RETRY:
return r
except httpx.TransportError:
pass
await asyncio.sleep(2 ** attempt + random.random())
return NoneWorkers and the queue
The crawl itself is a queue of URLs and a few workers taking from it. Each worker fetches a page, records it and queues every new link on the same site:
async def worker(client, queue, seen, robots, pages):
while True:
url = await queue.get()
try:
if len(pages) >= MAX_PAGES or not robots.can_fetch("*", url):
continue
r = await fetch(client, url)
if r is None or "text/html" not in r.headers.get("content-type", ""):
continue
pages.append((url, r.status_code))
parser = LinkParser()
parser.feed(r.text)
for href in parser.links:
link = urldefrag(urljoin(url, href)).url
if urlparse(link).netloc == urlparse(START).netloc and link not in seen:
seen.add(link)
queue.put_nowait(link)
finally:
queue.task_done()queue.task_done() in finally matters: the main function waits on queue.join(), which only returns once every queued URL has been marked done, skipped ones included.
robots.txt
Being a polite crawler means honouring the site's robots.txt. urllib.robotparser reads it; fetch it through the same client so it uses the proxy too:
from urllib.robotparser import RobotFileParser
robots = RobotFileParser()
r = await client.get(urljoin(START, "/robots.txt"))
robots.parse(r.text.splitlines() if r.status_code == 200 else [])An empty rule set (no robots.txt) allows everything.
The full crawler
import asyncio
import random
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin, urlparse
from urllib.robotparser import RobotFileParser
import httpx
PROXY = "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
START = "https://books.toscrape.com/"
MAX_PAGES = 200
WORKERS = 8
RETRY = {429, 500, 502, 503, 504}
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag == "a":
href = dict(attrs).get("href")
if href:
self.links.append(href)
async def fetch(client, url):
for attempt in range(4):
try:
r = await client.get(url)
if r.status_code not in RETRY:
return r
except httpx.TransportError:
pass
await asyncio.sleep(2 ** attempt + random.random())
return None
async def worker(client, queue, seen, robots, pages):
while True:
url = await queue.get()
try:
if len(pages) >= MAX_PAGES or not robots.can_fetch("*", url):
continue
r = await fetch(client, url)
if r is None or "text/html" not in r.headers.get("content-type", ""):
continue
pages.append((url, r.status_code))
parser = LinkParser()
parser.feed(r.text)
for href in parser.links:
link = urldefrag(urljoin(url, href)).url
if urlparse(link).netloc == urlparse(START).netloc and link not in seen:
seen.add(link)
queue.put_nowait(link)
finally:
queue.task_done()
async def main():
limits = httpx.Limits(max_connections=WORKERS, max_keepalive_connections=0)
headers = {"User-Agent": "Mozilla/5.0 (compatible; my-crawler/1.0)"}
async with httpx.AsyncClient(proxy=PROXY, limits=limits, timeout=30,
follow_redirects=True, headers=headers) as client:
robots = RobotFileParser()
r = await client.get(urljoin(START, "/robots.txt"))
robots.parse(r.text.splitlines() if r.status_code == 200 else [])
queue = asyncio.Queue()
seen, pages = {START}, []
queue.put_nowait(START)
tasks = [asyncio.create_task(worker(client, queue, seen, robots, pages)) for _ in range(WORKERS)]
await queue.join()
for t in tasks:
t.cancel()
print(f"crawled {len(pages)} pages")
for url, status in pages[:10]:
print(status, url)
asyncio.run(main())Run it with python crawler.py. Before pointing it at another site, read that site's terms, keep WORKERS low and add a delay if responses slow down.
Where to take it next
- Save what you find. Parse the fields you need from
r.text(withselectolax,lxmlorBeautifulSoup) and write them out as you go, instead of keeping pages in memory. - Target a country. Add
-country-us(or any code) to the username so the site sees local visitors; see the geo-targeting guide. - Pages behind a login. Switch those requests to a sticky session so they keep one IP: sticky vs rotating sessions.
- JavaScript-heavy pages. Render them with a headless browser: Playwright.
- Blocks and CAPTCHAs. Work through the debugging checklist.
Prefer a framework? Scrapy does the queue, retries and throttling for you.



