I have been studying on big data for a while. And i use, actually trying to use PySpark:). But in some point i really confused. For example as i know spark depending on its RDD option making parallelization automatically. And so why do we use clusters except using this local parallelization? Or do we use cluster mode for really big data(I am not talking about deploy mode i only say 2 or 3 or 4 slaves)? Actually i imagine parallelization like this, for example my computer have 12 cores so i think these 12 cores are individual computers and so like i have 12 computers. So because this thought it seems unnecessary to me to use a cluster for example in emr one master node and 2 slave nodes. And when i have 2 slaves is a parallelization keep going on them too. For example like i have 2 slaves and each of them 12 cores like my computer and so do i have 24 cores in this situation? If it is complicated and the title is wrong or deficient i can edit. Thanks in advance.