Back to problems

Concurrent Web Crawler

Object-Oriented Programming · Anthropic · Medium

This problem extends LeetCode 1242, Web Crawler Multithreaded; solving that earlier problem first is recommended. Design a concurrent web crawler that receives a starting URL startUrl and an HtmlParser, and returns every distinct reachable URL whose hostname is exactly identical to the hostname of startUrl. The crawler must obey these rules: Begin visiting from startUrl. A URL may be recorded and expanded only if its hostname exactly matches the hostname of startUrl. All…

Checking your access…