Do we need HDFS or S3 when running Spark on Kubernetes?
Will data locality be that efficient if we use just NFS storage type?
Or maybe there is something fundamentally wrong in my understanding of Spark on Kubernetes.
Do we need HDFS or S3 when running Spark on Kubernetes?
Will data locality be that efficient if we use just NFS storage type?
Or maybe there is something fundamentally wrong in my understanding of Spark on Kubernetes.
It depends. If you are working externally with data(HDFS/S3), then you won't have data locality and performance won't be awesome.
You can run hdfs inside Kubernetes to try and avoid this issue.