TL/DR: How to track and serve the input transformation for keras-flavored mlflow models?
Neural network training usually involves preprocessing steps in which
- continuous variables are scaled and shifted to have unit width and zero mean,
- categorical variables (integer or string) are transformed to one-hot encoding.
When the model is applied to new data, the scaling weights and the category-to-index association needs to be known.
In keras there are generally two options to perform preprocessing:
- Option 1: Using preprocessing layers, or
- Option 2: perform the transformation before the training when the dataset is loaded.
With Option 1, the transformation is part of the model and will be automatically applied when the network is used and served as a mlflow model.
My question concerns Option 2: What is the recommended way
- to keep track of the input transformation in mlflow for different experiments,
- and how to apply the same transformations when the model is served, e.g. with
mlflow model serve?