Differencies between MLflow deployment possibilities

Viewed 127

Can someone please explain what are the main use cases when deciding how to serve a model from MLflow:

  • using command line "mlflow models serve -m ...."
  • deploying local Docker container with the same model
  • deploying model online for example on AWS Sagemaker

I am mainly interested in differencies between option A and B because as I understand both can be accessed as REST API endpoints. And I assume if network rules are in place then both can be called also externally.

1 Answers

Imho, the main difference is described in the documentation:

NB: by default, the container will start nginx and gunicorn processes. If you don’t need the nginx process to be started (for instance if you deploy your container to Google Cloud Run), you can disable it via the DISABLE_NGINX environment variable

And the model serve uses only Flask, so it could be less scalable.

Related