How can I convert my model like mobilenet_v1_1.0_224_quant.tflite in tensorflow/examples?

Viewed 629

I am trying int8 quantization of my model on TensorFlow Lite. Conversion itself worked using tensorflow 1.15.3 but the converted model ran extremely slowly on Kirin 990. (Conversion using tensorflow 2.3.0 did not work.)

mobilenet_v1_1.0_224_quant.tflite in tensorflow/examples runs fast on Kirin 990.

So I checked the differences.

  1. My model is int8(tf.lite.OpsSet.TFLITE_BUILTINS_INT8) quantization but mobilenet_v1_1.0_224_quant.tflite seems uint8 quantization.

  2. The "filter" property of Conv2D has a "quantization" attribute in mobilenet_v1_1.0_224_quant.tflite, but The "filter" property of Conv2D has no "quantization" attribute in my converted model.

How can I convert my model like mobilenet_v1_1.0_224_quant.tflite?

2 Answers

I found a statement "UInt8 models trained with the legacy quantization-aware training path are also supported" in TensorFlow Lite Hexagon delegate. It means that you need quantization-aware training to get Uint8 weight. I think that this is the answer for my question though it is not useful since I need post-training quantization.

Related