CodingPhone, OnsiteSoftware EngineerReported Apr, 2026High Frequency
You receive:
startUrlHtmlParser interface that retrieves every URL linked from a specified pageBuild a crawler that returns every URL reachable from startUrl whose hostname is the same as startUrl. The result may be returned in any order.
The HtmlParser interface is defined below:
interface HtmlParser {
// Returns all URLs from a given page URL.
public List<String> getUrls(String url);
}
The function is invoked as follows:
List<String> crawl(String startUrl, HtmlParser htmlParser)
The crawler must:
startUrl.HtmlParser.getUrls(url) to retrieve the links on a page.startUrl.http protocol and as having no port.http://example.com/page#section1 should be treated. Should that URL differ from http://example.com/page#section2, or should both identify one URL? Discuss the choice with the interviewer if clarification is necessary.Once the single-threaded crawler is complete, create a multithreaded or concurrent version that can improve performance.
The concurrent crawler must:
Use a thread pool to cap the number of active threads.