Back to problems

Implement a high-throughput web crawler safely

System Design · Amazon · Hard

Build a concurrent web crawler, in pseudocode or a real language, that discovers pages in approximately breadth-first order while a separate worker pool continuously executes analysis tasks on pages that have already been downloaded. The crawler must enforce the following rules: Honor every site’s robots.txt and apply a configurable per-host request rate. Normalize and deduplicate URLs, including canonical redirects, so each distinct page is fetched and analyzed at most…

Checking your access…