Algorithm · Apple · Medium
Implement a compact TF-IDF ranking function. You are given a non-empty list of document strings docs and a query string q. Tokenize both the documents and the query using these rules: Convert all text to lowercase. Split at every character that is not a letter or digit. Discard any empty tokens produced by splitting. For a document d and token t, define: $$tf(t,d)=\text{number of occurrences of } t \text{ in } d$$ $$df(t)=\left \{ d \in docs : t \in d \}\right $$…
Checking your access…