- I am an extreme beginner and I like to write Python code using the Jupyter extension for VS Code.
- I recently started using Azure DataBricks, where my notebooks seem to be running on remote Spark clusters.
- I would like to be able to have the same experience as the Jupyter extension for VS Code, while running my notebooks on remote Spark clusters.
- Is there a way that I can do #3 with other types of Spark clusters? (Ideally I would like to do this with arbitrary Spark clusters.)
Question #4 depends on premises 1-3, which may be wrong or misleading.
I might not be explaining that correctly. I think I'm looking for something in VS Code analogous to what JupyterLab Integration does for JupyterLab:
JupyterLab Integration, on the other hand, keeps notebooks locally but runs all code on the remote cluster if a remote kernel is selected. This enables your local JupyterLab to run single node data science notebooks (using pandas, scikit-learn, etc.) on a remote environment maintained by Databricks or to run your deep learning code on a remote Databricks GPU machine.
But JupyterLab Integration seems to rely partially on SSH, so maybe there is a way to achieve part of this functionality outside of JupyterLab just by using SSH?