AzureML pass data between pipeline without saving it

Viewed 100

I have made two scripts using PythonScriptStep where data_prep.py prepares a dataset by doing some data transformation which is thereafter sent to train.py for training an ML model in AzureML.

It is possible passing data between pipeline steps using PipelineData and OutputFileDatasetConfig, however these seem to save the data in azure blob.

Q: How can I send the data between the steps without saving the data anywhere?

1 Answers

The data has to be passed somehow.

You can influence the storage account by changing the output datastore. If the data is just a collection of numbers, you can pass "dummy" data (e.g., empty text file) between the scripts, and have the upstream one log those numbers as metrics using Run.get_context().log(*) or MLFlow, and the downstream one load those values.

Fundamentally, there's no way to pass information between steps without it being stored somewhere, whether that's the "default blob store", another storage account, or metrics in the workspace.

Related