Guides4 min readOct 11, 2026

403, 429 or CAPTCHAs through a proxy? A debugging checklist

A proxy alone doesn't make a scraper look like a person. Work through this checklist to find out why a site is blocking you, from status codes and headers to pacing, sessions and choosing the right IP type.

By crawlproxies
403, 429 or CAPTCHAs through a proxy? A debugging checklist

You set up a proxy, the first requests work, and then the 403s start. Or every page is a CAPTCHA. A proxy changes where your requests come from, but sites look at much more than the IP: how fast you go, what your requests look like, and whether your behaviour adds up. Work through this list in order. Most blocks are solved in the first five steps.

1. Read the status code properly

CodeWhat it usually means
407 Proxy Authentication RequiredNot the website: your proxy credentials are wrong
403 ForbiddenThe site (or its firewall) decided not to serve you
429 Too Many RequestsYou're going too fast for this site
503 with a challenge pageA bot-protection check (often JavaScript-based)
200 but the page is a CAPTCHASoft block: you're suspected, not banned

A 407 is the only one you fix in your proxy settings. Everything else is about how the requests look and behave.

2. Confirm the proxy works

Before blaming the site, make sure traffic really goes through the proxy:

bash
curl -x http://USERNAME:PASSWORD@geo.crawlproxies.com:8080 https://ipinfo.io/json

If you see your own IP, the proxy isn't applied, and the site is blocking you.

3. Look like a real browser

Default HTTP clients announce themselves: python-requests/2.x, Go-http-client/1.1, curl/8.x. That alone gets you blocked on many sites. Send a realistic, complete set of headers:

python
HEADERS = {
    # Use a current browser version; an outdated one stands out too.
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
    "Accept-Encoding": "gzip, deflate",
}

(Only advertise br in Accept-Encoding if your client can decode Brotli; for requests that means pip install brotli.)

Keep the headers consistent with each other. A Windows Chrome user agent with Mac-only headers is a mismatch that sites notice.

4. Slow down

Rate limiting is the most common block of all, and no proxy fixes going too fast:

  • Limit concurrent requests per site, not just overall.
  • Add a short random pause between requests to the same site.
  • When you get a 429, honour the Retry-After header and back off exponentially.
python
import random, time, requests

def polite_get(url, proxies, tries=5):
    for attempt in range(tries):
        r = requests.get(url, headers=HEADERS, proxies=proxies, timeout=30)
        if r.status_code != 429:
            return r
        wait = float(r.headers.get("Retry-After", 2 ** attempt))
        time.sleep(wait + random.uniform(0, 1))
    return r

5. Match your location to the site

A US store visited by an IP from Vietnam with Accept-Language: de-DE is unusual. Target the country the site expects (USERNAME-country-us) and set Accept-Language to match. See the geo-targeting guide.

6. Keep sessions consistent

Logging in from one IP and loading the next page from another looks like a stolen session. For anything stateful (logins, carts, checkouts):

  • Use a sticky session (USERNAME-session-ID-time-1800) for the whole flow.
  • Keep cookies between requests (a requests.Session, or a browser context).
  • Don't share one session between unrelated accounts.

Sticky vs rotating sessions has patterns for this.

7. Use a real browser for JavaScript-heavy sites

Some protection runs JavaScript in the page and checks the result, and some looks at low-level details of how the client connects. A plain HTTP library can't pass those checks; a real browser does it naturally. If you see challenge pages that never resolve, switch to Playwright or Puppeteer, and block images and fonts to keep bandwidth down.

8. Choose the right proxy type

Datacenter IP ranges are public, so strict sites block them on sight. If you're on Datacenter and getting 403s from the first request, move to Residential, or Mobile for the strictest platforms. Choosing the right proxy compares them.

9. Rotate after a block, don't hammer

If an IP gets blocked, retrying on the same IP just confirms the block. On a rotating username a new connection gets a new IP; on a sticky session, switch to a new session id. Then slow down, because something in steps 3 to 6 is still giving you away.

10. Stay on the right side of the rules

Automated access has limits, technical and legal. Check a site's terms and robots.txt, keep your request rate reasonable, and be careful with personal data: privacy laws apply no matter which IP you send requests from. Sustainable scraping is polite scraping, and it gets blocked a lot less.

Still blocked?

Some sites simply refuse a whole category of traffic. If you've been through the list and still can't get through, open the chat bubble on this page and tell us the site and what you're seeing. We'll help you pick the right product and settings.

Written by
crawlproxies
Create account