Build a command-line utility that accepts a file path, reads the file, and writes to standard output every distinct line exactly once. This is a global deduplication — not just collapsing consecutive duplicates — so a line that appears multiple times anywhere in the file must be emitted only the first time it is encountered. The output must preserve the original first-appearance order of the unique lines.
The implementation must be self-contained and runnable end-to-end in your local development environment. A typical in-memory approach uses a hash set to remember which lines have already been seen, combined with a structure (such as a list or queue) to retain the order of first occurrences. The core challenge is not a complex algorithm but rather breaking the problem into clear steps and discussing each one with the interviewer.
Follow-up 1 — file too large for memory: the input file no longer fits in RAM. Describe an external-memory strategy: stream the file, hash each line, and partition the data by spilling to disk (for example, bucket lines by hash into multiple temporary files). Then deduplicate each bucket independently.
Follow-up 2 — preserve global order: explain how the external-memory design can still emit unique lines in the original first-seen order despite the partitioning step. Continue reasoning about how your chosen language handles file-system access and the loading of large inputs into memory.