I've been attempting to deploy a machine learning solution with Tensorflow Serving on an embedded device (Jetson Xavier [ARMv8])
One model used by the solution is a stock Xception network, generated by:
xception = tf.keras.applications.Xception(include_top=False, input_shape=(299, 299, 3), pooling=None
xception.save("./saved_xception_model/1", save_format="tf")
Running the Xception model on the device generates reasonable performance - about 0.1s to predict, ignoring all processioning:
xception = tf.keras.models.load_model("saved_xception_model/1", save_format="tf")
image = get_some_image() # image is numpy.ndarray
image.astype("float32")
image /= 255
image = cv2.resize(image, (299, 299))
# Tensorflow predict takes ~0.1s
xception.predict([image])
However, once the model is running in a Tensorflow Serving GPU container, via Nvidia-Docker, the model is much slower - about 3s to predict.
I've been trying to isolate the cause of the poor performance, and I've run out of ideas.
So far I've tested:
- Tweaking TF Serving's batching parameters to go all out on latency
(
batch_timeout_micros: 0,max_batch_size: 1, and noticed a modest 0.5s gain in performance. - Optimizing the model with TensorRT via
saved_model_cli. - Running the Xception model in isolation, as the only model being served by TF Serving.
- Experimenting with doubling the memory allocated per TF process.
- Experimenting with enabling and disabling batching altogether.
- Experimenting with enabling and disabling model warmup.
I would expect TF Serving to provide the same (more or less, allowing for GRPC encoding and decoding) prediction time as TF, and they do for other models I'm running. None of my efforts have got Xception upto the ~0.1s performance I would expect.
My install of Tensorflow is built by Nvidia, from TF version 2.0. My TF Serving container is self-built from the TF-Serving 2.0 source, with GPU support.
I start a Tensorflow Serving container as follows:
tf_serving_cmd = "docker run --runtime=nvidia -d"
tf_serving_cmd += " --name my-container"
tf_serving_cmd += " -p=8500:8500 -p=8501:8501"
tf_serving_cmd += " --mount=type=bind,source=/home/xception_model,target=/models/xception_model"
tf_serving_cmd += " --mount=type=bind,source=/home/model_config.pb,target=/models/model_config.pb"
tf_serving_cmd += " --mount=type=bind,source=/home/batching_config.pb,target=/models/batching_config.pb"
# Self built TF serving image for Jetson Xavier, ARMv8.
tf_serving_cmd += " ${MY_ORG}/serving"
# I have tried 0.5 with no performance difference.
# TF-Serving does not complain it wants more memory in either case.
tf_serving_cmd += " --per_process_gpu_memory_fraction:0.25"
tf_serving_cmd += " --model_config_file=/models/model_config.pb"
tf_serving_cmd += " --flush_filesystem_caches=true"
tf_serving_cmd += " --enable_model_warmup=true"
tf_serving_cmd += " --enable_batching=true"
tf_serving_cmd += " --batching_parameters_file=/models/batching_config.pb"
I'm at the point where I'm starting to wonder if this is a bug in TF-Serving, although I have no idea where (Yeah, I know it's never a bug, it's always the user...)
Can anyone suggest why TF-Serving might underperform compared to TF?