How to use Koalas in Dataproc from a local Jupyter Notebook

Viewed 138

Checking Google documentation, I was able to submit Spark jobs to a Dataproc cluster and install JupyterLab inside the cluster to run iterative operations on notebooks.

However, I could not discover the proper configuration to run iterative commands from a local Jupyer Notebook (on my machine) using DataProc cluster resources.

I'm especially interested in creating a cluster from my local JupyterLab and then using pySpark (Koalas) to perform a series of operations on large dataframes hosted on BigQuery and GCS. My target experience is to use Dataproc in my local JupyerLab in the same way it can be used accessing the JupyterLab installation inside the cluster machines or Vertex IA.

Does anyone know how to configure it?

1 Answers

To run local Jupyter notebook against a remote Dataproc cluster, you local machine needs to be able to connect to the cluster master node VM.

One way is to create a cluster with external IP and set up firewall rule to allow you to connect to the IP. But it is not secure, I don't recommend that.

Another way is to create an ssh tunnel from your local machine to the master node:

gcloud compute ssh ${HOSTNAME} \
    --project=${PROJECT} \
    --zone=${ZONE}  -- \
    -D ${PORT} -N

Then configure your local Spark/Jupyter to the local endpoint which will connect to the remote endpoint.

Related