FP32 Tensor to INT8 Quantization (with Asymmetric Quantization Follow-ups)

Given an FP32 tensor x (treat it as a 1-D array / contiguous buffer) and a floating-point scale, quantize it into an INT8 tensor q.

Base version (symmetric)

For each element:

  • `q[i] = (in...

Example

Unlock to view complete problem details

and practice with sample input/output

Was this article helpful?

View Test Cases & Run Code requires membership

Standard Input
Execution Result: