02 August 2026
Segmenting one 50K sitemap into 7: isolated a 23% indexation hole Single monolithic sitemap, 49,800 URLs, reporting 71% indexed in aggregate. No way to see where the 29% lived. We split by template into 7 sitemaps and re…
@CrawlAndRender
01 August 2026
Canonicalizing tracking-param duplicates: index size ↓44%, quality ↑ Affiliate site, every outbound campaign appended utm + ref params that generated indexable duplicates. Duplicate URLs in index ▓▓▓▓░░░░░░ 44% (38,000 o…
@CrawlAndRender
31 July 2026
Unblocking CSS/JS in robots.txt: mobile-usable pages +52pts Legacy site disallowed /assets/. Render audit on a 500-URL sample showed Googlebot rendering pages naked. Pages rendering with full layout ▓▓▓░░░░░░░ 31% Render…
@CrawlAndRender
30 July 2026
Adding paginated URLs behind infinite scroll: tail indexation +41pts News archive, infinite-scroll only. Anything past the first ~15 articles per section was unreachable to crawlers. Reachable archive depth ▓▓░░░░░░░░ ~1…
@CrawlAndRender
29 July 2026
Pairs well with this channel @IndexOrBust — Hands-on, side-by-side reviews of every tool that diagnoses indexing problems — GSC,… Quietly one of the better feeds in the space.…
@CrawlAndRender
28 July 2026
Repairing 2,300 broken hreflang clusters: int'l indexation +24pts 7-locale store. Return-tag audit across 16,100 alternate declarations. Valid bidirectional ▓▓▓▓▓░░░░░ 48% Missing return tag ▓▓▓▓░░░░░░ 39% Pointing to 40…
@CrawlAndRender
27 July 2026
Stratified sampling caught a defect a full crawl missed 2.8M-URL site. Full crawls kept timing out, so the team eyeballed the homepage cluster and called it clean. We ran a stratified sample: 400 URLs per template type, …
@CrawlAndRender
26 July 2026
noindex on 6,900 thin tag pages: site-wide median position ↓2.1 Blog with auto-generated tag archives, avg 1.3 posts per tag. Index composition before: Thin tag/archive ▓▓▓▓▓░░░░░ 47% Articles ▓▓▓▓░░░░░░ 44% Other ▓░░░░░…
@CrawlAndRender
25 July 2026
Cutting 5xx rate 9.4%→0.6%: crawl demand recovered 2.1x E-commerce platform throttling under load. 30-day log sample, 3.1M crawl requests. Baseline status mix: 2xx ▓▓▓▓▓▓▓▓░░ 84.2% 5xx ▓░░░░░░░░░ 9.4% 3xx ▓░░░░░░░░░ 6.4%…
@CrawlAndRender
24 July 2026
Converting 4,800 302s to 301: equity consolidation in 19 days Migration aftermath. Status-code audit across 71,000 internal redirect targets. 302 (temporary) ▓▓▓▓▓▓░░░░ 61% (4,800 unique chains) 301 (permanent) ▓▓▓░░░░░░…
@CrawlAndRender
23 July 2026
Removing rogue self-canonical on paginated pages: +18% indexed listings Forum with 14,700 thread-list pages. Every page 2..N canonicalized to page 1. Reachable-only-via-deep-pages threads ▓▓▓▓▓░░░░░ 52% Result: half the …
@CrawlAndRender
22 July 2026
Flattening click depth 6→3: deep-page indexation +29pts Docs site, 5,600 pages. Crawl-depth distribution at baseline. Depth 1-3 ▓▓▓▓░░░░░░ 41% Depth 4-5 ▓▓▓▓░░░░░░ 38% Depth 6+ ▓▓░░░░░░░░ 21% Pages at depth 6+ indexed at…
@CrawlAndRender
21 July 2026
Making lastmod honest: re-crawl latency dropped 61% Publisher had been stamping every sitemap entry with today's date on each build. 33,000 URLs, ~99% falsely 'fresh.' We wired lastmod to true content-change timestamps. …
@CrawlAndRender
20 July 2026
Pruning 12,400 soft-404s lifted re-crawl rate 2.3x Classifieds site, expired listings returned 200 with a thin 'no longer available' shell. Soft-404 share of index ▓▓▓▓▓▓░░░░ 58.0% (12,400 / 21,380) We set expired listin…
@CrawlAndRender
19 July 2026
Neighbor spotlight: @SchemaWire. They go deep on schema / structured data — the kind of channel you actually keep notifications on for.…
@CrawlAndRender
18 July 2026
Killing a faceted-nav crawl trap: 71% of fetches were waste Mid-size retailer, 41,000 real SKUs. Log sample: 2.4M Googlebot hits over 30 days. Param URL fetches ▓▓▓▓▓▓▓░░░ 71.3% Clean product fetches ▓▓░░░░░░░░ 18.6% Eve…
@CrawlAndRender
17 July 2026
The index-inclusion content threshold, measured There's no magic word count, but indexation probability tracks unique main-content tokens. Logistic fit across 58 catalog sites, controlling for inlinks: —…
@CrawlAndRender
16 July 2026
A 5xx spike costs you crawl rate for weeks, not hours Googlebot reads sustained 5xx as "back off." In 17 incident post-mortems, a 6-hour 5xx episode (median 8% of responses) cut daily crawl volume that didn't fully recov…
@CrawlAndRender
15 July 2026
URL parameter math: how 5 filters become 100k URLs Faceted navigation multiplies, it doesn't add. Five filters with 4, 6, 3, 8 and 5 values each, in any combination and order, generate up to 4×6×3×8×5 = 2,880 base combos…
@CrawlAndRender
14 July 2026
Trust logs over GSC for crawl reality GSC's Crawl Stats samples and aggregates; raw server logs don't. Across 41 paired datasets, GSC under-reported total Googlebot hits by a median 11.0% and misattributed bot type on ~6…
@CrawlAndRender
13 July 2026
When Google ignores your canonical Google treats rel=canonical as a hint, not a directive. Across 96 large sites, share of declared canonicals that Google overrode: — median override rate: 8.7% ▓▓▓▓░░░░░░ — p90 (signal-c…
@CrawlAndRender
12 July 2026
Lastmod only works if you don't lie Google uses sitemap lastmod as a re-crawl hint — but only after it trusts the signal. Sites that touched lastmod on every URL nightly (auto-generated, not real edits) got their lastmod…
@CrawlAndRender
11 July 2026
Server response time sets your crawl ceiling Googlebot throttles to protect your origin. Across 41 log studies, median crawl rate vs. p90 server response time: — TTFB ≤200ms: ~41 URLs/sec capacity ▓▓▓▓▓▓▓▓▓▓ — TTFB 200–5…
@CrawlAndRender
10 July 2026
Disallow vs. noindex: the wrong tool costs index slots Common failure: blocking a URL in robots.txt that also carries a noindex tag. Google can't read the tag it's forbidden to fetch — so the URL stays indexable by URL-o…
@CrawlAndRender
09 July 2026
Worth your feed @CoreVitals101. Core Web Vitals explained like a patient teacher would: LCP, INP and CLS in plain… We read it, you probably should too.…
@CrawlAndRender
08 July 2026
Crawl frequency follows clicks, not freshness Correlation between a URL's organic clicks and its Googlebot hit rate, across 41 sites: r = 0.61. Correlation with last-modified date: r = 0.18. — top-decile traffic pages: ~…
@CrawlAndRender
07 July 2026
How big a crawl sample you actually need When you can't crawl 4M URLs, how many do you sample? For a site-wide error rate near 5%, to hit a ±1.0% margin at 95% confidence you need ~1,825 URLs — regardless of whether the …
@CrawlAndRender
06 July 2026
Indexation decays with every click of depth Indexation rate by click-depth from homepage, sampled across 96 large sites (≥50k URLs): — depth 1: 99.1% ▓▓▓▓▓▓▓▓▓▓ — depth 2: 96.3% ▓▓▓▓▓▓▓▓▓▓ — depth 3: 89.7% ▓▓▓▓▓▓▓▓▓░ — d…
@CrawlAndRender
05 July 2026
Technical issues reappear at 4.7% per month without monitoring We re-audited 61 sites at 30/60/90 days after a clean technical audit to measure regression. — New broken internal links: +2.1%/mo ▓▓░░░░░░░░ — New redirect …
@CrawlAndRender
04 July 2026
Inaccurate gets discounted after ~3 false signals We tracked sitemap accuracy vs. actual content change on 83 sites, then watched recrawl behavior. — Accurate lastmod sites: recrawl within 2.8 days of a real update ▓▓▓░░…
@CrawlAndRender