[SYSTEM DESIGN] Web Crawler
PremiumDistributed web crawler at billion-page scale: BFS frontier, per-host politeness, Bloom-filter dedup, async fetcher pool, S3 + Kafka storage tier.
Что внутри
Web Crawler
The target is 1 billion successful page fetches per 30-day month: about 386/s average. A 1,200/s planned peak leaves room for retries and bursty hosts. With roughly one second of average in-flight network time, Little's Law gives about 1,200 concurrent requests at peak. Twenty pods with 250 in-flight slots each provide 5,000-request concurrency headroom; this is not the contradictory one-million-RPS fleet from the previous design.
Frontier, politeness and failure recovery
The frontier is a durable priority queue, not a global breadth-first queue. Tasks have leases/visibility deadlines and fencing tokens. A worker ACKs only after the durable completion contract; if it crashes, the task becomes visible again. Hostname ownership is an atomic lease shared by all secondary URL buckets, so splitting a hot host does not bypass politeness. Rate adapts to configured policy, latency, 429/503 responses and Retry-After.
RFC 9309 standardizes robots.txt parsing and caching but not Crawl-delay. The design may accept Crawl-delay as an optional vendor extension; it does not call it an RFC rule. Robots responses are cached under normal controls and generally not beyond 24 hours. When robots.txt is unreachable because of a server/network error, crawling is temporarily disallowed as RFC 9309 specifies.
DNS caches honor positive TTLs and RFC 2308 negative TTLs. Every redirect repeats URL normalization, DNS resolution, private/reserved-address rejection and robots policy. Fetches cap redirects, compressed and decoded bytes, decompression ratio and time to limit SSRF and decompression-bomb risk.
Dedupe and recrawl
Полный разбор, ADR-ы, сценарии и deep dives — после оплаты бандла.
System Design Cases
Полный доступ ко всем кейсам бандла
Premium открывает полный разбор для подготовки к интервью
- Где архитектура ломается первой и как защищать выбранный дизайн.
- Конкретный capacity math: размеры данных, throughput и пороги масштабирования.
- Trade-off-ы в стиле ADR, которые легко превращаются в структурированный ответ.
- Запускаемые сценарии: happy path, отказы, retry и recovery.
Регистрация бесплатна. Оплата — следующим шагом, из этого же кейса.
Уже есть аккаунт?