Desktop vs. cloud crawler: where the line actually sits
Benchmarked 312 audit runs by site size. Throughput crossover for a desktop crawler on 16GB RAM landed at ~487k URLs before swap thrash spiked run time.
— Under 50k URLs: desktop median 11 min ▓▓░░░░░░░░
— 50k–500k: desktop median 1h 48m ▓▓▓▓▓░░░░░
— Over 500k: desktop p90 timed out / OOM ▓▓▓▓▓▓▓▓▓▓
Cloud crawlers held a flat ~14 URLs/sec regardless of size but cost a median $0.0019/URL.
So what: desktop wins on cost and JS-render control below half a million URLs. Above it, RAM is the silent ceiling — list-mode chunking buys you ~2x, but past 700k the cloud's distributed fetch is cheaper than your time. Pick by URL count, not habit.
Crawl & Render
@CrawlAndRender
Desktop vs. cloud crawler: where the line actually sits
Этот пост опубликован в Telegram-канале Crawl & Render. Подписаться можно по ссылке: @CrawlAndRender.