I'm loading two models A and B using the same tensorflow model server instance (running inside a single Docker container). (using tensorflow_model_sever v2.5.1)
the models are ~5GB on disk, of which about 1.7GB is just the inference related nodes.
model A is only loaded once, and model B gets a new version every once in a while. my client only requests predictions from model A. model B isn't used. incidentally, both models have warmup data.
every time a new version is loaded for model B, the tail latency for graph runs in model A jumps from 20msec to > 100msec (!). I'm getting the tail latency both from tensorflow model server - :tensorflow:serving:request_latency_bucket, as well as from my GRPC client.
the container has plenty of available memory and CPU, this is also seen with a GPU.
I tried changing num_load_threads and num_unload_threads, flush_filesystem_caches, but to no avail.
so far didn't manage to get rid of it using GRPC hedging/manual double dispatch (but working on it), anybody ever seen this, and better yet managed to get over this? thanks!