Scrapy already knows how to use proxies: put a proxy URL in a request's meta and its built-in HttpProxyMiddleware does the rest, including the username and password. This guide wraps that in a small middleware, so every request gets a proxy automatically, then adds sticky sessions, geo-targeting and sensible crawl settings.
You'll need your crawlproxies username and password from the generator. The examples use the Residential gateway geo.crawlproxies.com:8080.
The one-line version
For a quick test, set the proxy on a request directly:
import scrapy
class IpSpider(scrapy.Spider):
name = "ip"
def start_requests(self):
yield scrapy.Request(
"https://ipinfo.io/json",
meta={"proxy": "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"},
)
def parse(self, response):
yield response.json()Scrapy reads the credentials from the URL and sends them as Proxy-Authorization for you.
A reusable middleware
Rather than touching every request, add a downloader middleware that sets the proxy for you. In middlewares.py:
from urllib.parse import quote
class CrawlproxiesMiddleware:
"""Route every request through crawlproxies.
request.meta options:
hqp_country = "us" exit country (two-letter code)
hqp_session = "cart42" keep the same IP for this session id
"""
def __init__(self, user, password, host, country):
self.user, self.password, self.host, self.country = user, password, host, country
@classmethod
def from_crawler(cls, crawler):
s = crawler.settings
return cls(s.get("HQP_USER"), s.get("HQP_PASSWORD"),
s.get("HQP_HOST", "geo.crawlproxies.com:8080"), s.get("HQP_COUNTRY"))
def process_request(self, request, spider):
if "proxy" in request.meta: # an explicit proxy wins
return None
user = self.user
country = request.meta.get("hqp_country", self.country)
if country:
user += f"-country-{country}"
session = request.meta.get("hqp_session")
if session:
user += f"-session-{session}-time-{request.meta.get('hqp_time', 1800)}"
request.meta["proxy"] = f"http://{quote(user)}:{quote(self.password)}@{self.host}"
return NoneEnable it in settings.py, before Scrapy's own proxy middleware (which runs at 750):
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.CrawlproxiesMiddleware": 350,
}
HQP_USER = "USERNAME"
HQP_PASSWORD = "PASSWORD"
HQP_COUNTRY = "us" # or None for worldwideKeep real credentials out of version control. Read them from the environment in settings.py with os.environ["HQP_PASSWORD"].
Sticky sessions for stateful crawls
Most crawls are happy with rotating IPs. When a spider has to log in or walk through a multi-step flow, pin those requests to one IP with a session id, and keep using the same id for the follow-up requests:
def start_requests(self):
yield scrapy.Request("https://example.com/login", callback=self.login,
meta={"hqp_session": "acct7"})
def login(self, response):
yield scrapy.FormRequest.from_response(
response, formdata={"user": "me", "pass": "secret"},
callback=self.after_login, meta={"hqp_session": "acct7"},
)Use a different session id per account. See sticky vs rotating sessions for patterns.
Making sure IPs really rotate
Scrapy reuses open connections for speed, and requests sent over the same proxy connection leave from the same IP. If you need every request to have its own IP, give each one a unique, short-lived session:
import uuid
def make_request(self, url):
return scrapy.Request(url, meta={"hqp_session": uuid.uuid4().hex[:10], "hqp_time": 60})Each unique session id gets its own exit IP, regardless of connection reuse.
Settings that keep you unblocked
A proxy hides where you are, not how fast you go. These settings matter as much as the IPs:
CONCURRENT_REQUESTS = 32
CONCURRENT_REQUESTS_PER_DOMAIN = 8
DOWNLOAD_DELAY = 0.5 # seconds between requests to the same site
RANDOMIZE_DOWNLOAD_DELAY = True
AUTOTHROTTLE_ENABLED = True # adapts the delay to how the site responds
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_TARGET_CONCURRENCY = 4
RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [429, 500, 502, 503, 504, 522, 524, 408]
DOWNLOAD_TIMEOUT = 30
ROBOTSTXT_OBEY = True
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"AutoThrottle slows down when the site slows down, which is exactly what a polite crawler should do.
Troubleshooting
| Symptom | Fix |
|---|---|
407 Proxy Authentication Required | Check HQP_USER / HQP_PASSWORD and the targeting syntax |
TunnelError | The proxy refused the tunnel: credentials, or the plan is out of bandwidth |
| Every response from the same IP | Connections are being reused: use unique short sessions as shown above |
Lots of 403 / 429 | Lower concurrency and add delay; see the debugging checklist |
| Middleware has no effect | Its priority must be below 750, and the module path must match your project |
Need a headless browser inside Scrapy for JavaScript-heavy pages? The same username format works in Playwright.



