How to ship and run spark-submit with virtualenv

Viewed 795

am trying to submit spark job on standalone cluster, I've zipped the virtualenv as venv.zip and I submit the job as shell script

#!/bin/sh
PYSPARK_PYTHON=./venv/bin/python \
PYSPARK_DRIVER_PYTHON=./venv/bin/python 
spark-submit \
--jars ojdbc6.jar \
--master spark://HOST:7077 \
--archives venv.zip#venv \
job.py

but I keep getting that modules are not found even though it exists in the venv and it runs fine in local mode.

I also tried to log into worker node and try run the venv, after activating the virtualenv manually , the modules can be found, it seems the scripts are using system-wide python, how can I fix this ?

1 Answers

I could do it with the below snippet, basically, I zipped the venv content and put the venv in HDFS (if you don't have HDFS or any shared accessible location by the nodes) if you don't have ... then I think you can clone the virtual envrionment on all nodes under same path

#!/bin/sh
PYSPARK_PYTHON=./venv/bin/python \
PYSPARK_DRIVER_PYTHON=./venv/bin/python \
spark-submit \
--conf spark.yarn.appMasterEnv.PYSPARK_PYTHON=./venv/bin/python \
--conf spark.executorEnv.PYSPARK_PYTHON=./venv/bin/python \
--conf spark.yarn.dist.archives=hdfs:///user/sw/python-envs/venv.zip#venv \
--conf spark.yarn.appMasterEnv.HADOOP_USER_NAME=hdfs \
--conf spark.executorEnv.HADOOP_USER_NAME=hdfs \
--master yarn \
--deploy-mode cluster \
--py-files p1.py,p2.py \
main.py
Related