Tutorials4 min readOct 11, 2026

Rotating proxies in Scrapy: a middleware that just works

Plug crawlproxies into Scrapy with a small downloader middleware. Rotating IPs by default, sticky sessions when a spider needs them, per-country targeting, and the settings that keep you from getting blocked.

By crawlproxies
Rotating proxies in Scrapy: a middleware that just works

Scrapy already knows how to use proxies: put a proxy URL in a request's meta and its built-in HttpProxyMiddleware does the rest, including the username and password. This guide wraps that in a small middleware, so every request gets a proxy automatically, then adds sticky sessions, geo-targeting and sensible crawl settings.

You'll need your crawlproxies username and password from the generator. The examples use the Residential gateway geo.crawlproxies.com:8080.

The one-line version

For a quick test, set the proxy on a request directly:

python
import scrapy

class IpSpider(scrapy.Spider):
    name = "ip"

    def start_requests(self):
        yield scrapy.Request(
            "https://ipinfo.io/json",
            meta={"proxy": "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"},
        )

    def parse(self, response):
        yield response.json()

Scrapy reads the credentials from the URL and sends them as Proxy-Authorization for you.

A reusable middleware

Rather than touching every request, add a downloader middleware that sets the proxy for you. In middlewares.py:

python
from urllib.parse import quote


class CrawlproxiesMiddleware:
    """Route every request through crawlproxies.

    request.meta options:
      hqp_country = "us"      exit country (two-letter code)
      hqp_session = "cart42"  keep the same IP for this session id
    """

    def __init__(self, user, password, host, country):
        self.user, self.password, self.host, self.country = user, password, host, country

    @classmethod
    def from_crawler(cls, crawler):
        s = crawler.settings
        return cls(s.get("HQP_USER"), s.get("HQP_PASSWORD"),
                   s.get("HQP_HOST", "geo.crawlproxies.com:8080"), s.get("HQP_COUNTRY"))

    def process_request(self, request, spider):
        if "proxy" in request.meta:          # an explicit proxy wins
            return None
        user = self.user
        country = request.meta.get("hqp_country", self.country)
        if country:
            user += f"-country-{country}"
        session = request.meta.get("hqp_session")
        if session:
            user += f"-session-{session}-time-{request.meta.get('hqp_time', 1800)}"
        request.meta["proxy"] = f"http://{quote(user)}:{quote(self.password)}@{self.host}"
        return None

Enable it in settings.py, before Scrapy's own proxy middleware (which runs at 750):

python
DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.CrawlproxiesMiddleware": 350,
}

HQP_USER = "USERNAME"
HQP_PASSWORD = "PASSWORD"
HQP_COUNTRY = "us"          # or None for worldwide

Keep real credentials out of version control. Read them from the environment in settings.py with os.environ["HQP_PASSWORD"].

Sticky sessions for stateful crawls

Most crawls are happy with rotating IPs. When a spider has to log in or walk through a multi-step flow, pin those requests to one IP with a session id, and keep using the same id for the follow-up requests:

python
def start_requests(self):
    yield scrapy.Request("https://example.com/login", callback=self.login,
                         meta={"hqp_session": "acct7"})

def login(self, response):
    yield scrapy.FormRequest.from_response(
        response, formdata={"user": "me", "pass": "secret"},
        callback=self.after_login, meta={"hqp_session": "acct7"},
    )

Use a different session id per account. See sticky vs rotating sessions for patterns.

Making sure IPs really rotate

Scrapy reuses open connections for speed, and requests sent over the same proxy connection leave from the same IP. If you need every request to have its own IP, give each one a unique, short-lived session:

python
import uuid

def make_request(self, url):
    return scrapy.Request(url, meta={"hqp_session": uuid.uuid4().hex[:10], "hqp_time": 60})

Each unique session id gets its own exit IP, regardless of connection reuse.

Settings that keep you unblocked

A proxy hides where you are, not how fast you go. These settings matter as much as the IPs:

python
CONCURRENT_REQUESTS = 32
CONCURRENT_REQUESTS_PER_DOMAIN = 8
DOWNLOAD_DELAY = 0.5              # seconds between requests to the same site
RANDOMIZE_DOWNLOAD_DELAY = True

AUTOTHROTTLE_ENABLED = True       # adapts the delay to how the site responds
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_TARGET_CONCURRENCY = 4

RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [429, 500, 502, 503, 504, 522, 524, 408]
DOWNLOAD_TIMEOUT = 30

ROBOTSTXT_OBEY = True
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"

AutoThrottle slows down when the site slows down, which is exactly what a polite crawler should do.

Troubleshooting

SymptomFix
407 Proxy Authentication RequiredCheck HQP_USER / HQP_PASSWORD and the targeting syntax
TunnelErrorThe proxy refused the tunnel: credentials, or the plan is out of bandwidth
Every response from the same IPConnections are being reused: use unique short sessions as shown above
Lots of 403 / 429Lower concurrency and add delay; see the debugging checklist
Middleware has no effectIts priority must be below 750, and the module path must match your project

Need a headless browser inside Scrapy for JavaScript-heavy pages? The same username format works in Playwright.

Written by
crawlproxies
Create account