Sagemaker not outputting Tensorboard logs to S3 during training

Viewed 916

I'm training a model with Tensorflow using Amazon Sagemaker, and I'd like to be able to monitor training progress while the job is running. During training however, no Tensorboard files are output to S3, only once the training job is completed are the files uploaded to S3. After training has completed, I can download the files and see that Tensorboard has been logging values correctly throughout training, despite only being updated in S3 once after training completes.

I'd like to know why Sagemaker isn't uploading the Tensorboard information to S3 throughout the training process?

Here is the code from my notebook on Sagemaker that kicks off the training job

import sagemaker
from sagemaker.tensorflow import TensorFlow
from sagemaker.debugger import DebuggerHookConfig, CollectionConfig, TensorBoardOutputConfig

import time

bucket = 'my-bucket'
output_prefix = 'training-jobs'
model_name = 'my-model'
dataset_name = 'my-dataset'
dataset_path = f's3://{bucket}/datasets/{dataset_name}'

output_path = f's3://{bucket}/{output_prefix}'
job_name = f'{model_name}-{dataset_name}-training-{time.strftime("%Y-%m-%d-%H-%M-%S", time.gmtime())}'
s3_checkpoint_path = f"{output_path}/{job_name}/checkpoints" # Checkpoints are updated live as expected
s3_tensorboard_path = f"{output_path}/{job_name}/tensorboard" # Tensorboard data isn't appearing here until the training job has completed

tensorboard_output_config = TensorBoardOutputConfig(
    s3_output_path=s3_tensorboard_path,
    container_local_output_path= '/opt/ml/output/tensorboard' # I have confirmed this is the unaltered path being provided to tf.summary.create_file_writer()
)

role = sagemaker.get_execution_role()

estimator = TensorFlow(entry_point='main.py', source_dir='./', role=role, max_run=60*60*24*5,
                           output_path=output_path,
                           checkpoint_s3_uri=s3_checkpoint_path,
                           tensorboard_output_config=tensorboard_output_config,
                           instance_count=1, instance_type='ml.g4dn.xlarge',
                           framework_version='2.3.1', py_version='py37', script_mode=True)

dpe_estimator.fit({'train': dataset_path}, wait=True, job_name=job_name)
2 Answers

There is a issue on tensorflow github related to the s3 client in version 2.3.1 which is the one you are using. Check in the cloudwatch logs if you have an error like

OP_REQUIRES failed at whole_file_read_ops.cc:116 : Failed precondition: AWS Credentials have not been set properly. Unable to access the specified S3 location

Then the provided solution is to add GetObejectVersion permission to the bucket. Alternatively, to confirm that is a tensorflow issue, you can try a different version.

First some speculation without any facts: Sagemaker could work as some other systems that sync files between local drive and s3. They might check that the file hasn't been accessed recently before syncing it so that they don't copy it while someone is writing to it. The log files are written constantly until shutdown so that might result in them not being copied ever.

I have used Sagemaker Docker containers with same problem. I've tried two ways circumvent this problem and they seemed to work.

First one is to periodically create a new log file. So e.g. every 30 minutes call again tf.summary.create_file_writer(...) to switch to a new log file. Old file is synced to s3 when it's not used anymore.

Second one is to directly write logs to s3. tf.summary.create_file_writer('s3://bucket/dir/'). This is more instant way of getting the info into s3.

Related