Preprocessing data for Sagemaker Inference Pipeline with Blazingtext

Viewed 848

I'm trying to figure out the best way to preprocess my input data for my inference endpoint for AWS Sagemaker. I'm using the BlazingText algorithm.

I'm not really sure the best way forward and I would be thankful for any pointers.

I currently train my model using a Jupyter notebook in Sagemaker and that works wonderfully, but the problem is that I use NLTK to clean my data (Swedish stopwords and stemming etc):

import nltk
nltk.download('punkt')
nltk.download('stopwords')

So the question is really, how do I get the same pre-processing logic to the inference endpoint ?

I have a couple of thoughts about how to proceed:

  • Build a docker container with the python libs & data installed with the sole purpose of pre-processing the data. Then use this container in the inference pipeline.

  • Supply the Python libs and Script to an existing container in the same way you can do for external lib an notebook

  • Build a custom fastText container with the libs I need and run it outside of Sagemaker.

  • Will probably work, but feels like a "hack": Build a Lambda function that has the proper Python libs&data installed and calls the Sagemaker Endpoint. I'm worried about cold start delays as the prediction traffic volume will be low.

I would like to go with the first option, but I'm struggling a bit to understand if there is a docker image that I could build from, and add my dependencies to, or if I need to build something from the ground up. For instance, would the image sagemaker-sparkml-serving:2.2 be a good candidate?

But maybe there is a better way all around?

0 Answers
Related