FP32 Tensor to INT8 Quantization (with Asymmetric Quantization Follow-ups)
Given an FP32 tensor x (treat it as a 1-D array / contiguous buffer) and a floating-point scale, quantize it into an INT8 tensor q.
Base version (symmetric)
For each element:
- `q[i] = (in...
Example
Unlock to view complete problem details
and practice with sample input/output
Was this article helpful?
View Test Cases & Run Code requires membership
Standard Input
Execution Result:
