Skip to content

When a site blocks you

How escalation works

Frankensurf starts with the cheapest tool and climbs a six-rung ladder only when a site pushes back.

Every way to fetch a page is a rung. Low rungs are fast and free; high rungs get through more. Each read starts low and climbs only as far as it has to.

Rung What it does Tools
T0 Markdown Asks the site for text/markdown HTTP
T1 Light fetch Plain HTTP or a reader, no browser HTTP, Scrapling, Jina Reader
T2 Browser A real browser renders the page Chromium, Steel, Crawl4AI, Browserbase, Hyperbrowser, Browserless, Kernel, Anchor, Cloudflare
T3 Stealth and signed Looks like a person, or signs as a verified agent Camoufox, Scrapling, Patchright, nodriver, signed requests
T4 Unblocker Paid services built for the hardest walls Firecrawl, ZenRows, fastCRW, Scrapfly, Bright Data, Zyte, Apify
T5 You A person clears the wall Human handoff

A tool “succeeding” isn’t enough. Frankensurf checks what came back and climbs when it sees:

  • an app shell or unrendered template (VISUAL_REQUIRED);
  • a rendered page with almost no text (EMPTY_PAGE);
  • a challenge or “Just a moment…” page (CAPTCHA, BLOCKED);
  • a sign-in page where content should be (AUTH_REQUIRED).

Sites also fake walls for bots: a sign-in redirect, or a 404, that a real browser never sees. So on a public read:

  • a NOT_FOUND from a plain fetch, or after another wall, gets one more try;
  • a sign-in wall stops the climb only after auth_wall_confirmations (default 5) tools in a row hit one;
  • after a suspected fake wall, the tools that look most different from the one that was fooled go next (fake_wall_confirmers: Jina Reader, Camoufox, then the paid unblockers).

A page that is plainly a site’s bot page (a final URL such as /help/bots.html or /captcha, or text like “thinks you are a bot”) counts as CAPTCHA, and a short page that only asks you to sign in, in English, Spanish, Portuguese, French, German, Italian or Dutch, counts as AUTH_REQUIRED.

Signed-in reads (identities and profiles) keep the strict rule: a sign-in wall there means the session needs renewing, and the climb stops.

A plain HTTP page can pass every check and still lack its content: a site that loads results with JavaScript serves only its header and footer as text. So when an automatic read lands on plain HTTP with scripts and under 5,000 characters of text (second_opinion_text_chars), Frankensurf also asks a browser and keeps whichever page has clearly more content (at least 1.5 times as much, and 1,000 characters more). receipt.second_opinion shows both. If the render fails or isn’t better, the HTTP page stands; the only cost is time.

Then every automatic read gets a structural check (completeness.py). It looks at the page’s shape:

  • Search and category pages (a query parameter, or a path such as /search) need at least 10 distinct same-site item links (product, listing or detail URLs; never assets or service and legal pages) or 4 prices.
  • Item pages need a price or 1,500 characters of text.
  • Other pages need 1,500 characters, or 400 without scripts.

An incomplete page is re-read along completeness_ladder: Scrapling, Camoufox, then Firecrawl, Zyte, Scrapfly and ZenRows when paid tools are allowed, then Patchright and Bright Data. Each step is an ordinary read pinned to one tool, with the usual checks and receipt. Frankensurf stops at the first complete page and keeps the most complete one it saw. A search page that passes only narrowly (fewer than completeness_borderline_items, 20, item links and few prices) also gets one read from the strongest allowed tool, a paid one when allowed. completeness_max_extra_reads (default 4) and completeness_deadline_seconds (default 150) bound the extra work, and receipt.completeness lists every step.

Free tools on the ladder race in pairs (completeness_parallel, default 2, up to 4). The first complete page wins and the other read is cancelled. Starts are half a second apart, so a site sees at most two extra reads at once. Paid tools always run one at a time, so racing never pays twice. On the 139-site held-out set this cut the free tier’s p90 from 50 s to 37 s and its median from 5.6 s to 4.5 s.

Two more shapes of a page that came back without its content are caught on every kind of page, not only search results:

  • Placeholders. Prices of $0 with no real price, or template values that leaked into the text ({{ price }}, NaN, undefined), mean the page had not loaded. completeness.placeholder is set and the read escalates.
  • Menus only. When more than 80% of the text is link text and there is no price, the read got the site’s navigation, not the page. completeness.link_text_share says how much.

Some pages are complete but hide their prices behind a choice (“select guests to see prices”). No stronger tool would show more, so these don’t escalate; completeness.needs_interaction quotes the page so your agent knows why there is no price.

Some wrong pages look right to every tool, so Frankensurf names them instead of climbing:

  • Off-query results. A site that ignores an unknown search parameter shows its default feed, which passes the structure check. When the query is known (from parameters such as q, kw or keywords, or from your expect_terms), a results page whose items and text never mention it comes back observed but marked completeness.off_query. receipt.next_step says the search URL is probably wrong. A page still loading its results escalates as before. Once you know the right URL, a site module template keeps it.
  • Pages a site module vouches for. When a saved module’s assertions pass, the page is complete even if it has few links, because the module knows its results live in JSON. A module’s invalid markers turn the site’s own error or empty-feed page into NOT_FOUND.
  • Site error pages. A redirect to an error or 404 path, an error title such as “Page not found”, or a short page saying the page doesn’t exist fails as NOT_FOUND. A plain fetch’s error page gets one confirming read from a different tool, since some sites serve fake 404s to bots; then the read stops.

A blocked site can take a dozen tools in turn before one gets through. Once an automatic read has run for 10 seconds (hedge_after_seconds, 0 turns it off), one more read starts beside it, pinned to the first free tool on the completeness ladder (Jina Reader by default). If that read brings back a complete page first, it wins and the climb is cancelled. A hedge page that is incomplete or failed never replaces the main read, and a hedge never uses a paid tool. When the hedge wins, receipt.hedge says so and the site’s route hint learns it, so the next read starts there. On the 139-site held-out set, reads that used to take about 50 seconds (Michaels, Naukri, Anthropologie, Urban Outfitters) came back in 8 to 17.

Three things reorder the ladder so repeat visits are fast:

What it does Setting
Escalation After 4 walls in one read, allowed paid tools move ahead of the free ones left. escalate_after_walls
Site hints When a read succeeds only after earlier tools failed, the tool that got through goes first on that site for 24 hours. origin_route_hint_ttl_seconds
Route memory After 3 clean reads in a row on the same path pattern, that tool goes first for that path. route_memory_min_samples

Measured on g2.com: 63.5 s before hints, 8.5 s on the first read with escalation, 2.4 s with the hint.

A few sites need a particular tool from the first visit. Those live as data in bundled_route_seeds.json, never as site code, and your own settings always win.

Frankensurf is polite by default.

  • Reads to one site are at least 2 seconds apart, across processes (origin_min_interval_seconds).
  • A BLOCKED, CAPTCHA or RATE_LIMITED answer pauses direct reads of that site for 15 minutes (origin_cooldown_seconds).
  • Paid unblockers and handoff aren’t held by the pause: a different service or a person is handling the wall.
You want Set
One specific tool, no fallback provider="zenrows"
An exact order provider_candidates=["http", "camoufox"]
Paid tools allowed, within a budget allow_paid_fallbacks=True, max_cost_usd=0.05
Nothing on your machine allow_local_browser=False