Algorithm · Moveworks · Medium
Given two strings a and b, compute the Jaccard similarity of their token sets. Tokenize each string independently: Convert all letters to lowercase. Treat any character that is not an English letter as a separator, including spaces, punctuation, and digits. Discard all empty pieces left between separators. Build a set of unique tokens, so repeated words from the same string are counted once. Let A be the token set from a and B be the token set from b. The similarity is…
Checking your access…