I am inferencing model using tflite with GPUDelegateV2 from tf2.X, but the initial time seems to be so long, check the table as below:
The model I'm using comes from here, and I test these data using benchmark_model following here
TF version: 2.0.0 benchark_model parameters:
- use_gpu=true
- num_threads=1
- allow_fp16=true
- others by default
|----------------------|------------------|------------------|
|v3-small_224_1.0_float|inference time(ms)| initial time(ms) |
|----------------------|------------------|------------------|
| on CPU | 46.245 | 2.988 |
| on GPUDelegateV2 | 13.015 | 4480.14 |
|----------------------|------------------|------------------|
| mobilenet_v2_0.75_224|inference time(ms)| initial time(ms) |
|----------------------|------------------|------------------|
| on CPU | 123.116 | 2.849 |
| on GPUDelegateV2 | 27.084 | 6358.39 |
I tried with different models / devices but got results just alike (faster initial on newer cores but still much longer than how they work on CPU).
Is there anyone have some idea about this?