Is RDD kept in memory or is it immediately flushed out of memory after an action is completed?

Viewed 194

I am going through a book, which to me has stated a contradictory statement. Quoting the book: "Spark’s RDDs are by default recomputed each time you run an action on them." But in the next few lines, it states: "After computing it the first time, Spark will store the RDD contents in memory and reuse them in future actions."

My question is, if RDDs are stored in memory, why is it recomputed each time an action is called on them?

In the first statement it says RDD is recomputed each time, and in the second statement it says, RDD is stored in memory to reuse them for future actions.

1 Answers

"Spark’s RDDs are by default recomputed each time you run an action on them." For your this statement, yes RDDs are recomputed each time you run an action on them. Now the reason behind it is that if it store all the RDDs content in memory then your memory would get exhausted very soon. So, it cannot keep each and every RDDs in memory. When you perform any action on it, it reads from the source data and performs transformations on it and give you the output for your action.

"After computing it the first time, Spark will store the RDD contents in memory and reuse them in future actions." It does not store it in memory by default but based on your use case you can persist the specific RDD using df.cache() or df.persist() and then it will store that RDD content in memory and when you perform any action second time of the RDD that depends upon the cached RDD then it won't read that from the source but it would use it from memory. You only should cache the RDD if you are performing multiple actions on it or there is a complex transformation logic that you don't want spark to perform every time an action is called.

Related