You are asked to design a production service that converts uploaded PDF files into Markdown. You do not need to implement the raw processing primitives; instead, compose them into a scalable system.
Assume these building blocks already exist:
split_pdf(pdf) -> List[np.ndarray]: CPU-heavy; splits a PDF into individual page images.
ocr(page_array) -> str (or batched ocr(batch) -> List[str]): GPU-heavy; extracts text and layout from page images.
to_markdown(ocr_output) -> str: memory-heavy; converts OCR output into Markdown and/or assembles pages into a final document.
Together these create a three-stage pipeline in which different stages pressure different resources: CPU, GPU, and memory. The service must allow those stages to scale independently.
Working constraints:
- Document sizes range from a single page to a few thousand pages. The synchronous path may receive documents around 1,000 pages, while the asynchronous path must handle many concurrent jobs under bursty traffic.
- CPU splitting, GPU OCR, and memory-heavy Markdown conversion each create a different resource bottleneck.
- GPU capacity is the main cost driver; OCR throughput depends on GPU utilization and batch size, and idle GPUs are expensive.
- Final Markdown output must preserve original PDF page order regardless of the order in which pages complete.
- A fully rasterized 1,000-page PDF is too large to keep in one process's memory at once.
- Synchronous requests must finish or stream within a bounded connection lifetime, typically measured in minutes rather than hours.
Before designing, resolve these points with the interviewer if needed:
- Distribution of PDF sizes, including median and p99 page counts, plus expected request rate and concurrency.
- Whether the synchronous response must be one fully ordered document, or whether the API may stream completed pages incrementally.
- Latency SLOs for interactive synchronous requests versus delayed asynchronous results.
- Multi-tenancy and fairness requirements, including quotas and isolation.
- Durability, retention, and compliance constraints for uploaded PDFs and generated Markdown.
- Whether partial results or progress reporting are required, and whether at-least-once delivery with idempotency is acceptable.
Part 1: Synchronous API for one very large document
A single caller submits one very large PDF, roughly 1,000 pages, and expects the complete Markdown result as quickly as possible over one request/response interaction. Optimize end-to-end latency for this single document.
Your design must specify:
- The synchronous API contract and delivery model: one ordered final response body versus incrementally streamed output, with a clear rationale.
- A page-level decomposition that overlaps CPU splitting, GPU OCR, and memory-heavy Markdown processing in a pipeline instead of running the stages serially across all pages; explain why this overlap is the primary latency improvement.
- An intra-document GPU batching scheme, such as micro-batching pages belonging to that one document, including the throughput-vs-latency tradeoff and how GPU idle time is avoided.
- A result-ordering mechanism that reassembles out-of-order page completions into correct page order.
- A strategy for documents that exceed the connection lifetime ceiling, including what the client observes.
Part 2: Asynchronous API for many concurrent requests
Many clients submit conversion jobs concurrently, and each client is willing to receive results later through polling, webhooks, or download links. Optimize global throughput and resource utilization rather than the latency of any single job.
Your design must specify:
- An asynchronous API contract: submission returns a job identifier, plus status/progress and result retrieval; submission must be idempotent.
- Durable queues between the three pipeline stages to absorb bursts, decouple scaling, and let each stage autoscale on its own signal.
- Cross-job GPU batching that fills OCR batches from a shared queue across many jobs, allowing small jobs to share GPU capacity and keeping the GPU saturated.
- Fairness and scheduling controls so one large document or tenant cannot starve small jobs, such as fair queuing, per-tenant limits, or optional priority classes.
- Independent autoscaling for CPU, GPU, and Markdown worker pools based on their respective bottleneck metrics.
Be prepared to discuss:
- If a synchronous 1,000-page conversion would exceed the configured connection ceiling, does the service reject the request, degrade gracefully, or transparently fall back to asynchronous processing? What does the client see?
- How can the system guarantee fairness so one tenant submitting many huge documents cannot monopolize the GPU pool and starve other tenants?
- If OCR quality is poor on certain pages because of skew or low DPI, how would you support selectively re-processing only those pages with different settings, such as re-rasterizing at higher DPI, deskewing, or using a different OCR model, without reprocessing the entire document?
- If GPU OCR becomes the dominant operating cost, which levers would you consider, such as batch size, dynamic batching, quantization, queue-depth-based autoscaling, and spot or preemptible GPUs, and what risks do they introduce?