Scrapy docs mirror-keeper: bloom+sha256 CDC crawler, live-proven, shipped as skill
DONE 2026-07-03. crawlers/ package on SitemapSpider: md-first per URL (probed live: platform.claude.com serves real .md → source:md; claude.com/blog|customers|connectors|skills|plugins + anthropic.com/news + support 404 the variant → deterministic stdlib HTML distiller fallback). Skip logic exactly as requested: persisted deterministic bloom filter (sha256-derived k-hashes, 512KB bitset, pure stdlib vs awesome-scrapy Redis variants) = cross-run URL skip; sha256 manifest = write skip (unchanged pages never rewrite, git stays clean). PROVEN live: 10 blog pages written → rerun bloom-skipped all 10 with zero fetches → recrawl=1 re-fetched and SKIP-SHA d all 10 → platform target wrote 5 real-markdown docs. Typed TARGETS table (new mirror = one dict entry). Two real bugs found+fixed live: 404 .md responses never reached parse (handle_httpstatus_list) and the bloom filtered its own sitemap request on rerun (discovery layer now exempt — new posts stay discoverable). Skill+command: /scrapy-docs-crawler in cwc-engineering. ALSO: disk-full incident #2 mid-build (Chrome cache 4.3GB 1h after cask install; Desktop Commander proved to be the zero-disk break-glass shell) — logged in incident-runbook. scrapy 2.13.4 recorded in durable-mac-toolchain (7→8 entries).
created 2026-07-03 17:55:47 · updated 2026-07-03 17:55:47