I am using TF 1.15.0 and running the model in inference mode. For 10000 short texts, most of the time, it is using gpu around 20%, which seems normal since the inference is done for each example (batch_size=1). However, when the processing is close to the end, I noticed that GPU usage is getting less and less, and eventually, it stopped using GPU at all.
My program used to work normally and it took around 12 minutes to process the 10000 texts, but now it takes around 19 minutes to complete it.
I have never encountered this phenomena. What might cause this? It's a bert model for a classification task.