I use Sagemaker's SKLearnProcessor.run for executing my training job. Between the time my processing job starts executing and the time my first line of the code in the processing.py file is read, there is a delay of 4-5 minutes. After the job starts executing, irrespective of how large the input file is, the job completes execution quickly, as is expected from Sagemaker's processing capabilities.
My question is, can I somehow reduce the time it takes to start executing my processing.py file.
sklearn_job.run(code= os.path.join('s3://',bucket, code_prefix, 'preprocessing_v2.py'),
'''
inputs=[ProcessingInput(
input_name='raw1',
source= os.path.join('s3://',bucket, input_prefix, 'file1.csv'),
destination='/opt/ml/processing/input1'),
ProcessingInput(
input_name='raw2',
source= os.path.join('s3://',bucket, input_prefix, 'file2.csv'),
destination='/opt/ml/processing/input2')],
outputs=[ProcessingOutput(output_name='sample_file',
source='/opt/ml/processing/dataset',
destination=os.path.join('s3://',bucket, output_prefix))],
arguments=["--train_size", "0.8","--test_size","0.2"],
wait=True, logs=True,
)
'''