I have long running job in DataProc Cluster which Query Kafka Topic, I am using GCP Bucket as checkpointing, in case of any failure or job stop how I ensure that I start the job from same checkpoint i.e read from same data where the failure/stopping occur
Currently in code I have defined this as :
.option("checkpointLocation", "gs://checkpoint"+ uuid);
So in case of any failure how I can restart a job from same checkpoint ? And also I need to submit a new Job right in case of failure ?