High availability and geographic redundancy for Dataflow

Viewed 651

What is the best architecture in terms of HA for Dataflow on Google Cloud? My workloads are running in two regions. The Dataflow reads from one multi-regional bucket and writes out results into another multi-regional bucket.

To achieve HA (in case one of the regions becomes unavailable), I am planning to run two identical Dataflow pipelines, one in each separate region.

The question is whether this is viable architecture, especially in terms of writing results to the same multi-regional buckets. Pipeline uses TextIO which overrides files if they exist. Do you envision potential problems with that?

Thank you!

2 Answers

As long as GCP Dataflow spreads the workers in a zonal GCE instances within the same particular region, managed as a MIG groups, with any disaster across the location zone will require the user to restart the job and specify the zone in the separate region.

Given said this, we might assume that Dataflow offers a zonal high availability model rather then regional one, therefore by now it's not feasible to specify multiple regions and have Dataflow automatically failover to a different region in case of computational zone outage.

In the mentioned use case, I assume that for a Dataflow batch job which doesn't consume any real-time arriving data, you can just re-run this job at any time without data loss in case of failure. If the aim stays on ingesting data continuously discovering fresh files appearance in the GCS bucket, then probably you would need to launch streaming execution for this particular pipeline.

I would recommend you to look at Google Cloud Functions, that gives you an opportunity to compose the user function triggering the specific action based on some cloud event occurrence. I guess this way you might be able to fetch up the harmful event for the batch Dataflow pipeline in the prime regional zone and based on this execute then the same job in a separate computation region.

It would be even more beneficial for the community to file a feature request to the vendor via issue tracker considering Dataflow multi-region high availability implementation.

For your architecture, it seems that the question hinges on whether it is OK to have two TextIO transforms writing the same data to the same location.

No, it is probably not OK.

The exact flow of elements through your pipeline is not deterministic. Therefore the output of the two pipelines will not necessarily be byte-for-byte identical. They may even result in different numbers of file shards, depending on how you configure TextIO. So, especially in a failure situation, you could end up in a situation where some shards are from one pipeline, some are from the other pipeline, or even inconsistent numbers of shards. (you'll see some files named like 0000-of-0250 while others named like 0000-of-0242).

I have not reviewed the code to determine exactly the failure mode. TextIO does do some work to write everything to a temporary location, checkpoint, then move them to their final destination and checkpoint again. But I do not believe it is robust to your proposed use.

Related