How do you use ADF PipelineParameters with ML studio to change output location instead of input location

Viewed 127

I want to parameterise my ML studio pipeline such that it outputs to a different blob store depending on which ADF environment it is being run from - the dev or prod ADF instance. This is so that the data engineers can have an output on their dev blob storage so they can avoid developing in live, while still using the exact same ML Studio pipeline as would be used in live.

I have followed the instructions in this notebook to change the input data source on blob dynamically using PipelineParameters.

Can I / how do I do the same for output locations?

Below is the code I used. I have used ./outputs as the output folder. I want to change this to a location in blob storage dynamically set via a PipelineParameter that can be called from ADF. (Or another way if you have other suggestions!)

Thanks in advance for any suggestions!


from azureml.data.datapath import DataPath, DataPathComputeBinding
from azureml.pipeline.core import PipelineParameter

default_data_path = DataPath(
    datastore=Datastore(workspace, "prod_data_science"), 
    path_on_datastore='path/to/data/')

input_datapath_pipeline_param = PipelineParameter(name="input_datapath", default_value=default_data_path)

input_data_consumption = (
    input_datapath_pipeline_param, 
    DataPathComputeBinding(
        mode='mount',
        overwrite = True
    )
)

toy_step = PythonScriptStep(
    script_name="toy_code.py",
    source_directory="./",
    arguments=[
        '--input_folder', input_data_consumption,
        '--output_folder', './outputs',
    ],
    inputs=[input_data_consumption],
    runconfig=runconfig,
)
0 Answers
Related