Back to problems

Stream Deduplication with Near-Duplicate Detection

Algorithm · Perplexity · Medium

You are given a sequence of strings arriving one after another. Process the sequence in order and output a stream that contains only the first occurrence of each string under exact matching rules. In a follow‑up, extend the idea to near‑duplicate detection. Two strings are near‑duplicates when their edit (Levenshtein) distance is below a supplied threshold. Normalization (lowercasing and punctuation removal) can optionally be applied before the distance computation so that…

Checking your access…