Can you create new directories in azure ml designer

Viewed 428

So I'm creating a web service using the azure machine learning designer, I can safely import libraries from the Script Bundle and load files from it. But I can't seem to create new directories in it and save files inside this new directory. Tried using os.mkdirs and pathlib but didn't manage to do it. Is there any way to do so? Or in the current version this is not supported yet

Edit: So to reproduce the problem i'm facing wrote two python scripts connected to a test Script Bundle in the Azure designer, now this zip file only contains a test txt file to make sure the script bundle isn't empty. As follows in the images the first script creates a new directory in the bundle and creates a DataFrame containing the directories in the root of the zip file, the next script is responsible for creating and saving a txt file inside this newly created directory. What follows are screenshots of the problem and the scripts used:

How the pipline structure looks like

The first py script creating the new dir

Trying to save a txt file to this new dir

1 Answers

Each Execute Python Script module is a different temporary container. The folder you've created in one Execute Python Script only exists on that specific context, and the new Execute Python Script can't see it. This is the same for installing libraries for instance. If you need a library in more than one module, you have to pip-install it again.

The alternative to this is using the AML SDK to persist the folder and files in a datastore.

import pandas as pd
from azureml.core import Run, Datastore

def azureml_main(dataframe1 = None, dataframe2 = None):

    file_path = 'new_file_test.txt'
    file = open(file_path, mode = 'w')
    file.write('xxx') 
    file.close()
 
    run = Run.get_context(allow_offline=True)
    ws = run.experiment.workspace
    datastore = ws.get_default_datastore()
    datastore.upload_files(files = [file_path], target_path = "new_file_test.txt", overwrite = True)

    return dataframe1,

In a different module, you can download the file to the module execution context and use it:

import pandas as pd
from azureml.core import Run, Datastore

def azureml_main(dataframe1 = None, dataframe2 = None):

    run = Run.get_context(allow_offline=True)
    ws = run.experiment.workspace
    datastore = ws.get_default_datastore()
    
    datastore.download('new_file_test.txt', prefix='/new_file_test.txt/')

    return dataframe1,

Check the documentation of Execute Python Script, AzureBlobDatastore.upload_files(), and AzureBlobDatastore.download().

Related