should I pre-install cran r packages on worker nodes when using sparkr

Viewed 1878

I want to use r packages on cran such as forecast etc with sparkr and meet following two problems.

  1. Should I pre-install all those packages on worker nodes? But when I read the source code of spark this file, it seems that spark will automatically zip packages and distribute them to the workers via --jars or --packages. What should I do to make the dependencies available on workers?

  2. Suppose I need to use functions provided by forecast in a map transformation, how should I import the package. Do I need to do something like following, import the package in the map function, will it make multiple import: SparkR:::map(rdd, function(x){ library(forecast) then do other staffs })

Update:

After reading more source code, it seems that, I can use includePackage to include packages on worker nodes according to this file. So now the problem becomes is it right that I have to pre-install the packages on nodes manually? And if that's true, what's the use case for --jars and --packages described in question 1? If that's wrong, how to use --jars and --packages to install the packages?

3 Answers

a better choice is to pass your local R package by spark-submit archive option, which means you do not need install R package in each worker and do not install and compile R package while running SparkR::dapply for time consuming waiting. for example:

Sys.setenv("SPARKR_SUBMIT_ARGS"="--master yarn-client --num-executors 40 --executor-cores 10 --executor-memory 8G --driver-memory 512M --jars /usr/lib/hadoop/lib/hadoop-lzo-0.4.15-cdh5.11.1.jar --files /etc/hive/conf/hive-site.xml --archives /your_R_packages/3.5.zip --files xgboost.model sparkr-shell")

when call SparkR::dapply function, let it call .libPaths("./3.5.zip/3.5") first. And you need notice that the server version R version must be equal your zip file R version.

Related