Recovery of data stored on disk by Spark workers

Viewed 8

Spark stores shuffle data on disk irrespective of whether we call persist on it. Shuffles are expensive so that's understandable. This usually goes in spark.local.dir

What I want to know is, if a worker restarts/replaced and if there was some data stored in spark.local.dir, then, when the worker comes back up, does the newly spawned spark worker be able to reuse the contents from the directory?

The storage can be network attached like EBS for the case of worker getting replaced

0 Answers
Related