Is it possible to launch Spark application from a host with no Spark installed on it

Viewed 952

I have a remote host set up with Spark standalone instance (one master and one slave on the same machine for now). I also have local Java code with spark-core dependency and a packaged jar with actual Spark Application. I'm trying to start it using SparkLauncher class as described in it's Javadoc.

Here is dependency:

        <groupId>org.apache.spark</groupId>
        <artifactId>spark-core_2.10</artifactId>
        <version>${spark.version}</version>

And here is the code of the louncher:

        new SparkLauncher()
            .setVerbose(true)
            .setDeployMode("cluster")
            .setSparkHome("/opt/spark/current").setAppResource(Resources.getResource("validation.jar").getPath())
            .setMainClass("com.blah.SparkTestApplication")
            .setMaster("spark://"  + sparkMasterHostWithPort))
            .startApplication();

The error I'm getting is either path not found /opt/spark/current/ or, if I remove setSparkHome call, Spark home not found; set it explicitly or use the SPARK_HOME environment variable.

Here is my naive question(s): is there any workaround allowing me not to have Spark binaries installed on the local host where I want to run only the Launcher? Why Spark Java code referenced in the dependencies is not capable / is not enough to connect to some configured remote Spark Master and submitting the application jar? Even if I put Spark binaries, application code and if needed even the Spark Java jar to hdfs location and use other deployment approach, like YARN, would it be enough to use Launcher just to trigger submission and start remotely?

The reason is that I want to avoid installing Spark binaries on multiple client nodes only to submit and start dynamically created/modified Spark applications from there, it sounds like a waste to me. Not to mention necessity to package application in jar for each submission.

1 Answers

Short answer: you must have spark binaries on the client machine and SPARK_HOME environment variable pointing to it.

Long answer: however if you want to launch the job on remote cluster then you could make use of the following configurations in your spark job:

val spark = SparkSession.builder.master("yarn") 
.config("spark.submit.deployMode", "cluster")
.config("spark.driver.host", "remote.spark.driver.host.on.the.cluster") 
.config("spark.driver.port", "35000")
.config("spark.blockManager.port", "36000") 
.getOrCreate()

spark.driver.port and spark.blockManager.port are not mandatory, but needed if you are working in a closed environment, like let's say kubernetes network, and have some port gateway service defined for spark client pod.

Having remote host defined in master setting of the SparkLauncher will not work. You need to get the hadoop configurations from the cluster, usually it is located in /etc/hadoop/conf on the cluster nodes. Place hadoop config directory in the client machine and point HADOOP_CONF_DIR environment variable to it. This should be enough to get started.

Related