Suppose you are working in a Jupyter notebook with a labeled binary text classification dataset for separating messages into two categories: spam and not_spam.
Your implementation should do the following:
spam), or macro/micro averaging.Keep these requirements in mind:
Any correct output showing the requested information is accepted; the exact F1 value and predictions depend on your model and data split.
Example 1:
Input:
train_texts = ["Limited offer, claim your prize now", "Can we move our sync to 3pm?", "Weight loss pills online", "I'll review the document tonight"]
train_labels = ["spam", "not_spam", "spam", "not_spam"]
test_texts = ["You have been selected for a reward", "Please share the agenda before the call"]
Output:
F1 score (positive class: spam): 0.67
Predicted labels: ["spam", "not_spam"]
Explanation: The first test message is classified as spam, while the second is classified as not_spam.
Example 2:
Input:
train_texts = ["URGENT: update your payment details", "Lunch plans still on?", "Free money guaranteed", "The report is attached"]
train_labels = ["spam", "not_spam", "spam", "not_spam"]
test_texts = ["Verify your account immediately", "Thanks for sending the notes"]
Output:
F1 score (positive class: spam): 0.50
Predicted labels: ["spam", "not_spam"]
Predicted probabilities (spam): [0.92, 0.18]
Explanation: The displayed probabilities are optional and show the model's confidence that each test message belongs to the spam class.
Constraints:
spam or not_spam.