Crawl a list of pages, read each one's title, download every image it references, and report what happened — concurrently, and without one bad page sinking the run.
python -m pytest tests -q # 41 tests, 26 failing
python main.py # crawls the bundled demo site
Python 3.11+, standard library only. Third-party HTTP and scraping libraries are
out of scope — asyncio, html.parser and urllib.parse are what you have.
Nothing needs network access: the suite serves its own site on a local port.
45 minutes.
Gates: tests/test_parser.py, tests/test_fetcher.py
src/parser.py and src/fetcher.py are written and run. There are three
defects: two in src/parser.py, one in src/fetcher.py. Each contradicts the
docstring directly above it.
All three look fine on the happy path. A page whose <title> sits on one line,
and an <img> whose src is already an absolute URL, both come out correct —
which is why they shipped.
Gate: tests/test_crawl.py
crawl and _crawl_one in src/crawler.py are stubs. The docstring carries
the whole contract:
{"url", "title", "images_downloaded", "error"}<img> is a fact about the site, not about the crawlcrawl never raises for a per-URL failureGate: tests/test_performance.py
This is what the task exists for, and where a working part 2 can still be the wrong answer. Two budgets, over 40 pages with 3 images each:
| gate | what it measures | budget |
|---|---|---|
test_pages_are_crawled_concurrently | wall clock for the whole crawl | 3.0 s |
test_the_limit_is_global_not_per_page | peak simultaneous connections | ≤ concurrency |
Both are properties of how the work is scheduled, not of the answer. A sequential crawler and an unbounded one both return byte-identical results — same rows, same order, same counts — so no correctness test can tell them apart. That is why there are two of these, and why neither substitutes for the other.
The fixture server counts connections itself, so the second gate measures what your crawler really did rather than what it claims.
src/
fetcher.py async HTTP over asyncio streams (part 1 fix)
parser.py title and <img> extraction (part 1 fixes)
storage.py where images land (correct as shipped)
crawler.py the crawl (part 2 stub, part 3 budgets)
main.py run it; with no arguments, against the bundled demo site
tests/
fixture_site.py the site the suite crawls
tests/ is unchanged.concurrency, shrinking the fixture, or skipping
images.