System Design · Atlassian · Hard
Requirements Allow clients to provide one or multiple starting URLs for an image-crawling task. The platform must visit each page and its linked pages, collect image URLs, and record every image under its originating top-level URL. Clients need a way to retrieve a job's state, for example in_progress or completed. Clients must be able to request the image URLs discovered for any URL they submitted. The architecture should support an unbounded number of root URLs and no fixed…
Checking your access…