Performance penalty for Tensorflow Serving in docker hosted app

Viewed 674

I am experiencing a large performance penalty for calls to Tensorflow Serving, when the calling app is hosted in a docker container. I was hoping someone would have some suggestions. Details below on the setup and what I have tried.

Scenario 1:

  • Docker (version 18.09.0, build 4d60db4) hosted Tensorflow model, following the instructions here.
  • Flask app running on the host machine (not in container).
  • Using gRPC for sending the request to the model.
  • Performance: 0.0061 seconds / per prediction

Scenario 2:

  • Same docker container hosted Tensorflow model.
  • Container hosted Flask app running on the host machine (inside same container as the model).
  • Using gRPC for sending the request to the model.
  • Performance: 0.0107 / per prediction

In other words, when the app is hosted in the same container as the model, performance is ~40% lower.

I have logged timing on nearly every step in the app and have tracked the difference down to this line:

result = self.stub.Predict(self.request, 60.0)

In the container hosted app, the average round-trip for this is 0.006 seconds. For the same app hosted outside the container, the round-trip for this line is 0.002 seconds.

This is the function I am using to establish the connection to the model.

def TFServerConnection():
    channel = implementations.insecure_channel('127.0.0.1', 8500)
    stub = prediction_service_pb2.beta_create_PredictionService_stub(channel)
    request = predict_pb2.PredictRequest()
    return (channel, stub, request)

I have tried hosting the app and model in separate containers, building a custom Tensorflow Serving containers (optimized for my VM), and using the REST api (which decreased performance slightly for both scenarios).

Edit 1

To add a bit more information, I am running the docker container with the following command:

docker run \
    --detach \
    --publish 8000:8000 \
    --publish 8500:8500 \
    --publish 8501:8501 \
    --name tfserver \
    --mount type=bind,source=/home/jason/models,target=/models \
    --mount type=bind,source=/home/jason/myapp/serve/tfserve.conf,target=/config/tfserve.conf \
    --network host \
    jason/myapp:latest

Edit 2

I have now tracked this down to being an issue with stub.Predict(request, 60.0) in Flask apps only. It seems Docker is not the issue. Here are the versions of Flask and Tensorflow I am currently running.

$ sudo pip3 freeze | grep Flask
Flask==1.0.2

$ sudo pip3 freeze | grep tensor
tensorboard==1.12.0
tensorflow==1.12.0
tensorflow-serving-api==1.12.0

I am using gunicorn as my WSGI server:

gunicorn --preload --config config/gunicorn.conf app:app

And the contents of config/gunicorn.conf:

bind = "0.0.0.0:8000"
workers = 3
timeout = 60
worker_class = 'gevent'
worker_connections = 1000

Edit 3

I have now narrowed the issue down to Flask. I ran the Flask app directly with app.run() and got the same performance as when using gunicorn. What could Flask be doing that would slow the call to Tensorflow?

0 Answers
Related