CloudWatch only capturing last metric produced by SageMaker training job

Viewed 288

I am using Amazon SageMaker to train a PyTorch model and attempting to visualise the loss values in CloudWatch. I create my estimator:

from sagemaker.pytorch import PyTorch

estimator = PyTorch(
    entry_point="train.py",
    source_dir=source_dir,
    role=role,
    framework_version=framework_version,
    py_version="py3",
    train_instance_count=1,
    train_instance_type=instance_type,
    hyperparameters=hyperparameters,
    metric_definitions=[
        {"Name": "train:loss", "Regex": "Train Loss:([0-9\\.]+)"},
        {"Name": "val:loss", "Regex": "Val Loss:([0-9\\.]+)"},
    ],
    enable_sagemaker_metrics=True
)

and execute the training job:

estimator.fit(s3_url)

Which runs successfully, but when I look at the algorithm metrics in CloudWatch for the training job this creates, it seems to only capture the last reported loss values. This is also the case when using the TrainingJobAnalytics:

from sagemaker.analytics import TrainingJobAnalytics

analysis = TrainingJobAnalytics(training_job_name=estimator._current_job_name)
df = analysis.dataframe()
df

where the output looks like:

    timestamp   metric_name value
0         0.0   train:loss  0.471061
1         0.0   val:loss    0.167700

In the CloudWatch logs there are multiple values being logged but they do not appear to be captured. I was wondering whether someone could provide some advice for how to fix this?

0 Answers
Related