MLflow: Why can't backend-store-uri be an s3 location?

Viewed 768

I'm new to mlflow and I can't figure out why the artifact store can't be the same as the backend store?

The only reason I can think of is to be able to query the experiments with SQL syntax... but since we can interact with the runs using mlflow ui I just don't understand why all artifacts and parameters can't go to a same location (which is what happens when using local storage).

Can anyone shed some light on this?

1 Answers

MLflow's Artifacts are typically ML models, i.e. relatively large binary files. On the other hand, run data are typically a couple of floats.

In the end it is not a question of what is possible or not (many things are possible if you put enough effort into it), but rather to follow good practices:

  • storing large binary artifacts in an SQL database is possible but is bound the degrade the performance of the database sooner or later, and this in turn will degrade your user experience.
  • storing a couple of floats from a SQL database for quick retrieval for display in a front-end or via command line is a robust industry-proven classic

It remains true that the documentation of MLflow on the architecture design rationale could be improved (as of 2020)

Related