TensorFlow Serving vs. TensorFlow Inference (container type for SageMaker model)

Viewed 1223

I am fairly new to TensorFlow (and SageMaker) and am stuck in the process of deploying a SageMaker endpoint. I have just recently succeeded in creating a Saved Model type model, which is currently being used to service a sample endpoint (the model was created externally). However, when I checked the image I am using for the endpoint, it says '.../tensorflow-inference', which is not the direction I want to go in because I want to use a SageMaker TensorFlow serving container (I followed tutorials from the official TensorFlow serving GitHub repo-using sample models, and they are deployed correcting using the TensorFlow serving framework).

Am I encountering this issue because my Saved Model does not have the correct 'serving' tag? I have not checked my tag sets yet but wanted to know if this would be the core reason to the problem. Also, most importantly, what are the differences between the two container types-I think having a better understanding of these two concepts would show me why I am unable to produce the correct image.


This is how I deployed the sample endpoint:

model = Model(model_data =...)

predictor = model.deploy(initial_instance_count=...)

When I run the code, I get a model, an endpoint configuration, and an endpoint. I got the container type by clicking on model details within the AWS SageMaker console.

2 Answers

There are two APIs for deploying TensorFlow models: tensorflow.Model and tensorflow.serving.Model. It isn't clear from the code-snippet which one you're using, but the SageMaker docs recommend the latter for deploying from pre-existing s3 artifacts:

from sagemaker.tensorflow.serving import Model

model = Model(model_data='s3://mybucket/model.tar.gz', role='MySageMakerRole')

predictor = model.deploy(initial_instance_count=1, instance_type='ml.c5.xlarge')

Reference: https://github.com/aws/sagemaker-python-sdk/blob/c919e4dee3a00243f0b736af93fb156d17b04796/src/sagemaker/tensorflow/deploying_tensorflow_serving.rst#deploying-directly-from-model-artifacts

it says '.../tensorflow-inference', which is not the direction I want to go in because I want to use a SageMaker TensorFlow serving container

If you haven't specified an image argument for tensorflow.Model, SageMaker should be using the default TensorFlow serving image (seems like "../tensorflow-inference").

image (str) – A Docker image URI (default: None). If not specified, a default image for TensorFlow Serving will be used.

If all of this seems needlessly complex to you, I'm working on a platform that makes this set up a single line of code -- I'd love for you to try it, dm me at https://twitter.com/yoavz_.

There are different versions for the framework containers. Since the framework version I'm using is 1.15, the image I got had to be in a tensorflow-inference container. If I used versions <= 1.13, then I would get sagemaker-tensorflow-serving images. The two aren't the same, but there's no 'correct' container type.

Related