In Data Health, when should I use build status vs transaction status vs job status

Viewed 66

Data Health can check the status of the most recent build and/or of the most recent transaction. Under what conditions would one of these checks fail while the other passed? In other words, why would it matter which check I picked, and would I ever want both?

1 Answers

Sync status checks check whether the most recent sync of the dataset to another database succeeded or failed. It's actually not for Magritte syncs (which are effectively Foundry Builds of the dataset they produce), but for syncs of the dataset to other places, like Phonograph or Postgres.

Job status isn't a rename of transaction status. Transaction status was removed some time ago (unfortunately I don't know the context why), and job status was added. There is some amount of overlap but it's not a simple rename.

When to use which between Build Status and Job status:

Use a Build status check when the dataset is an output of a build and you want to check that the whole build, on all datasets including this dataset succeeded. Use a Job status check when the dataset is an intermediate dataset of the build, and you want to check whether the dataset got updated, regardless of whether other datasets in the build were successfully updated.

Build status and job status will be equivalent if the dataset is the only output of a build. They may differ if the dataset is an intermediate dataset or if the build has multiple outputs, and the job on the dataset succeeds (or doesn't run) but other jobs in the build fail, causing the build to fail.

Regarding "If your upstream dataset fails, and a downstream target dataset fails, will the downstream target set with a 1) job status health check pick this up 2) but a build status health check would not?": It depends on the build graph but generally the opposite is true: If a failure happens upstream, the build will fail but the job may never run, so it will neither succeed nor fail and the check won't catch this:

Build status applies to builds which final dataset is the dataset you're looking at. Transaction status applies to any builds that contains your dataset. This mirrors the way Job Tracker can show you either builds which final dataset is the one you're looking at, or any builds that touched it (the top-right toogle in Job Tracker).

If you are building this dataset directly, there is no difference. If you're building the dataset as part of a bigger pipeline, then Build status will never apply, while Transaction status might fail/succeed.

Most of the times, you probably care about Transaction status.

Note: build status check should be configured to the end dataset of any schedule otherwise it will ping in health checks over and over again if a intermediate dataset fails. Ex. If it is configured on intermediate dataset and that dataset fails, it will ping consistently over and over despite the end dataset building successfully. Therefore you either configure a "JOB" status check as oppose to a build or you configure the build status check on the end dataset.

workflow

Related