I am using Amazon SageMaker to train a PyTorch model and attempting to visualise the loss values in CloudWatch. I create my estimator:
from sagemaker.pytorch import PyTorch
estimator = PyTorch(
entry_point="train.py",
source_dir=source_dir,
role=role,
framework_version=framework_version,
py_version="py3",
train_instance_count=1,
train_instance_type=instance_type,
hyperparameters=hyperparameters,
metric_definitions=[
{"Name": "train:loss", "Regex": "Train Loss:([0-9\\.]+)"},
{"Name": "val:loss", "Regex": "Val Loss:([0-9\\.]+)"},
],
enable_sagemaker_metrics=True
)
and execute the training job:
estimator.fit(s3_url)
Which runs successfully, but when I look at the algorithm metrics in CloudWatch for the training job this creates, it seems to only capture the last reported loss values. This is also the case when using the TrainingJobAnalytics:
from sagemaker.analytics import TrainingJobAnalytics
analysis = TrainingJobAnalytics(training_job_name=estimator._current_job_name)
df = analysis.dataframe()
df
where the output looks like:
timestamp metric_name value
0 0.0 train:loss 0.471061
1 0.0 val:loss 0.167700
In the CloudWatch logs there are multiple values being logged but they do not appear to be captured. I was wondering whether someone could provide some advice for how to fix this?