Back to problems

Implement Top-p (Nucleus) Sampling in NumPy

Algorithm · Scale AI · Medium

Given the next-token probability vector produced by a language model, write a NumPy-based implementation of top-p (nucleus) sampling. Task Create sample_top_p(probs, p, rng) that draws and returns a token id from probs according to this procedure: Order token ids from highest probability to lowest probability. Select the shortest leading group whose accumulated probability is at least p (with 0 = 0, and sum(probs) = 1. p: A floating-point nucleus cutoff satisfying 0 = 1.…

Checking your access…