Back to problems

Implement TF–IDF with sparse matrices

Object-Oriented Programming · Thumbtack · Hard

Design and implement a TF–IDF vectorizer from scratch. You are given a collection of documents, where each document is a string. The implementation must: Build a memory-efficient tokenizer that converts text to lowercase and removes punctuation. Support optional document-frequency filtering through min_df and max_df. Compute term frequencies within each document. Use the smoothed inverse document frequency formula $$\text{idf}_t = \log\left(\frac{1 + N}{1 + df_t}\right) +…

Checking your access…