Figma · ML & AI Fundamentals
Clarify a BERT Keyphrase Span-Extraction Task
TrueInterview
October 7, 2026 · 2 min read
Clarifying a BERT Keyphrase Span-Extraction Task
The archived report points to a BERT keyphrase span-extraction coding exercise, but the linked statement itself is not preserved. As a result, the callable signature, provided model outputs, tokenizer behavior, valid spans, scoring rule, and required return value are all unknown. Describe how you would first reconstruct a deterministic contract, then pick an approach only after those details are confirmed, and finally validate the extraction pipeline.
Constraints & Assumptions
- The only technical fact that survives is that the exercise concerns BERT keyphrase span extraction.
- It is not known whether the exercise provides text, tokens, token labels, boundary scores, candidate spans, or a trained model.
- BIO tagging, overlap rules, confidence aggregation, output ordering, and training requirements are not part of the preserved record.
- Any implementation you propose must depend on an explicitly confirmed contract for input, output, and evaluation.
Clarifying Questions to Ask
- What exactly is passed to the function, and which portion of the pipeline is the candidate expected to implement?
- Which tokenizer and offset mapping are authoritative, particularly for subwords and normalized text?
- What qualifies as a valid span, and can keyphrases overlap, nest, repeat, or contain punctuation?
- What exact values and order must be returned, and how should invalid input and tied scores be handled?
Part 1 — Recover the contract
List the minimum interface, model-output, span-validity, and result rules required to make the exercise deterministic.
What This Part Should Cover
- Input types, who owns tokenization, and the character-offset convention
- Provided model outputs and where the candidate's implementation begins and ends
- Valid spans, how scores are interpreted, and any overlap policy
- Exact return representation, ordering, tie handling, and invalid-input behavior
Part 2 — Select a conditional approach
Compare appropriate approaches for at least two plausible confirmed contracts, without asserting that either one was the original exercise.
What This Part Should Cover
- A token-state scan when a full label sequence is provided
- Candidate construction and constrained selection when boundary or span scores are provided
- How overlap, maximum length, ordering, and score semantics affect the algorithm
- Complexity expressed in terms of the confirmed input size
Part 3 — Validate the pipeline
Describe tests that separately exercise tokenization, model-output interpretation, span construction, and final offset reconstruction.
What This Part Should Cover
- Empty input, punctuation, Unicode, repeated phrases, and subword boundaries
- Malformed or inconsistent label and offset inputs
- Overlapping, nested, touching, and tied-score candidates when the contract allows them
- Exact span-level examples and evaluation aligned to the confirmed output format
Hint — Pin down the implementation boundary first: A model name and task label do not tell you whether you are implementing tokenization, decoding provided outputs, ranking spans, or running a trained model.
What a Strong Answer Covers
- A source-faithful account of what is known and what is missing
- A deterministic contract established before any algorithm is chosen
- Conditional approaches that follow from the confirmed representations and rules
- Offset handling, edge cases, tests, and a candid complexity analysis
Follow-up Questions
- How would the design change if the input were BIO labels instead of boundary scores?
- Which tests expose an offset error caused by subword tokenization or text normalization?
- How would you evaluate exact spans separately from matches that only partially overlap?
Overview: Clarify a BERT keyphrase span-extraction task by recovering the interface, tokenizer and offset rules, valid spans, outputs, tests, and conditional implementation options.