A notebook-based ML coding screen. You are given a movie review, IMDB style, and must output the key span:
the start and end token positions of the fragment that best sums up the reviewer's verdict. The fragment need not
be a full sentence. The model is encoder-style span extraction with BERT (a BertModel plus a linear head that
scores every token as a start and as an end), not generative prompting.
notebook.ipynb: the notebook you complete. Configuration, data loading, the epoch loop and the final report are
written; the cells marked YOUR CODE are yours. A markdown cell above each one states the exact interface.span_utils.py: provided helpers for loading the data, mapping the character span to token positions through
the tokenizer's offset mapping, building loaders, and evaluation.data/: 3,000 training reviews (10% become a validation split) and 400 held-out test reviews, each with an
annotated key span given as character offsets; plus the WordPiece vocabulary for the tokenizer.check_notebook.py: runs the notebook without Jupyter and re-checks the result.Everything runs offline. The default encoder is a real transformers.BertModel built from a small BertConfig
(random initialisation) with BertTokenizerFast, standing in for bert-base-uncased; set BERT_MODEL to a
checkpoint name to run the same notebook with pretrained weights.
KeyphraseExtractor.__init__(bert_model_name, bert_config=None) builds the encoder
and a linear head with two outputs per token.forward(input_ids, attention_mask) returns (start_logits, end_logits), each (B, L).decode_spans(start_logits, end_logits, attention_mask, max_span_len) and
KeyphraseExtractor.predict return one (start, end) per review: the highest start + end score with
start <= end, at most max_span_len tokens, never on [CLS], [SEP] or padding.span_loss, the cross-entropy of the start logits against the gold start index and of the end logits
against the gold end index, averaged.train_one_epoch(model, loader, optimizer, scheduler), the training step (zero_grad, forward, loss,
backward, optimizer and scheduler step).The notebook must then run top to bottom and train to at least 90% accuracy on the test reviews, where accuracy is start-position accuracy and end-position accuracy averaged. Exact match and token-overlap F1 are reported next to it.
python check_notebook.py executes the notebook and trusts none of its printed numbers. It checks that the model
holds one BertModel and one Linear(hidden, 2) head, that forward equals the head applied to
last_hidden_state, that decode_spans matches a brute-force search on random logits (including cases where the
best start comes after the best end, or where [CLS] and padding score highest), that span_loss equals the
averaged cross-entropies and reaches both heads, that the training loss fell, and it re-scores the trained model on
the test file itself (accuracy of at least 0.90). Both follow-up answers must be written out.