I'm using tensorflow-gpu 2.5.0 for an object detection model. In the model, I'm calling the prediction function using model(x). When I run the testing in batch, I found that repeated calling the prediction on a same image will result in a faster inference. This gives an inaccurate inference time estimation for my model. What is the reason behind the speed improvement?
Example below is the inference time result from a repeated inference on three different images. It was executed in a loop where the format is <image number> - <inference time>. The inference speed become faster after the model seen image 1 and image 2 once. When I add a new image 3, the first time the model took longer to predict and subsequently become faster as well.
1 - 3.5671939849853516 seconds
2 - 1.1461808681488037 seconds
1 - 0.07942032814025879 seconds
2 - 0.08655834197998047 seconds
1 - 0.0813601016998291 seconds
2 - 0.08380460739135742 seconds
1 - 0.07466459274291992 seconds
2 - 0.08526778221130371 seconds
3 - 1.0617506504058838 seconds
3 - 0.07965445518493652 seconds