I've deployed a custom model with an async endpoint. I want to process video files with it because videos can have ~5-10 minutes I can't load all frames to memory. Of course, I want to make an inference on each frame.
I've written
input_fn - download video file from s3 using boto and creates generator which loads video frames with a given batch size - return a generator - written with OpenCV
predict_fn - iterate over generator batched frames and generate prediction using model - save prediction in list
output_fn - transform prediction into json format, gzip all to reduce the size
Endpoint works well, but the problem is concurrency. The sagemaker endpoint processes request after request (from cloudwatch and s3 save file time). I don't know why this happens. max_concurrent_invocations_per_instance is set to 1000. Other settings from PyTorch serving are as follows:
SAGEMAKER_MODEL_SERVER_TIMEOUT: 100000
SAGEMAKER_TS_MAX_BATCH_DELAY: 10000
SAGEMAKER_TS_BATCH_SIZE: 1000
SAGEMAKER_TS_MAX_WORKERS: 4
SAGEMAKER_TS_RESPONSE_TIMEOUT: 100000
And still, it doesn't work. So how can I create an async inference endpoint with PyTorch to get concurrency?