Thumbtack · ML & AI Fundamentals
Detail NLP preprocessing and n‑gram choices
TrueInterview
October 7, 2026 · 1 min read
Walk me through your text preprocessing pipeline for each input modality: typed text, scanned or handwritten OCR, and speech-to-text output. Include your choices for language detection and handling, normalization steps (case, punctuation, Unicode), tokenization method (plain whitespace, rule-based, or subword approaches such as BPE/WordPiece), stopword removal, lemmatization versus stemming, treatment of emojis, URLs, and code, and how you deal with out-of-vocabulary terms. You employed n-gram sizes from 1 to 3: give both theoretical and empirical reasons for these choices—address sparsity, vocabulary size, context span, and how they affect linear models as opposed to tree-based or neural models; show how performance and feature importance shifted between unigram, bigram, and trigram settings. Compare word-level and character-level n-grams and say when each is useful (for misspellings, morphology). Finally, describe how you would validate the pipeline (train/validation split, leakage checks) and how this setup compares with a modern transformer-based tokenizer and embedding.
Overview: This question probes a data scientist's understanding of NLP preprocessing and feature construction, including modality-specific text normalization, tokenization and subword decisions, n-gram range and sparsity trade-offs, management of OOV terms, emojis, URLs, and code, and empirical validation and model comparison.