Back to problems

File Deduplication

Algorithm · Anthropic · Easy

Suppose you are examining a set of files stored on disk. For each file you have its absolute path (a string with no whitespace) and its exact content (also a whitespace-free string). Two or more files are deemed duplicates when their contents match character-for-character. Your job: gather every cluster of duplicate files. A cluster must contain every path that shares the same content. Only clusters of size 2 or larger should be reported. Input: The first line provides an…

Checking your access…