I am currently using tensorflow models of yolov3 & yolov4 for inference (tf 2.2) As i needed to improve performance of my models, I was thinking of quantization, and going post training from fp32 model to fp16 All ressources I've found about that topics are only using tflite. Tflite is not an option, because my model will run on GPU and dockers. Is there any ressource about quantization without tflite?