Back to problems

Distributed Training Data Pipeline (FAR)

System Design · Amazon · Hard

Requirements Ingest raw text from object storage (S3), with Kafka supported as an optional streaming input. Build the flow as: extraction → tokenization → deduplication (both exact matching and MinHash/LSH-based near-duplicate removal) → quality screening (heuristics followed by a classifier) → sharded outputs. The system must operate at petabyte scale, tolerate idempotent reruns, assign shards deterministically, and expose observability for every stage. Petabyte-scale data…

Checking your access…