I have an ONNX, I have some doubts:
1, If I don't consider prediction accuracy, what is the fastest inference time and how can I obtain it ?
2, I found that, FP8 quantization is not much faster than INT8 quantization, sometimes even slower.
3, If I choose int8+fp16 as computational precision, use the default int8-max quant config setting is not get good performance, so do I still need to manual adjust the config based on the info of the engine by trex ?
4, Set autotune: bool = True in modelopt.onnx.quantization.quantize() get worse performance than manual setting ?
Thanks !
I have an ONNX, I have some doubts:
1, If I don't consider prediction accuracy, what is the fastest inference time and how can I obtain it ?
2, I found that, FP8 quantization is not much faster than INT8 quantization, sometimes even slower.
3, If I choose
int8+fp16as computational precision, use the default int8-max quant config setting is not get good performance, so do I still need to manual adjust the config based on the info of the engine by trex ?4, Set
autotune: bool = Truein modelopt.onnx.quantization.quantize() get worse performance than manual setting ?Thanks !