Back to problems

FP32 Tensor to Int8 Quantization

Algorithm · NVIDIA · Medium

Requirements Write the following function: The initial version must do the following: Accept an FP32 tensor x with any shape. Scale each element by dividing it by scale. Round or convert the results so they fit in the signed INT8 range. Produce an int8 tensor whose shape matches x. Follow-ups Extend the function with a zero_point argument to support asymmetric quantization. Describe situations in which symmetric quantization does not provide a good representation. Derive how…

Checking your access…