I want to use r packages on cran such as forecast etc with sparkr and meet following two problems.
Should I pre-install all those packages on worker nodes? But when I read the source code of spark this file, it seems that spark will automatically zip packages and distribute them to the workers via --jars or --packages. What should I do to make the dependencies available on workers?
Suppose I need to use functions provided by
forecastin amaptransformation, how should I import the package. Do I need to do something like following, import the package in the map function, will it make multiple import:SparkR:::map(rdd, function(x){ library(forecast) then do other staffs })
Update:
After reading more source code, it seems that, I can use includePackage to include packages on worker nodes according to this file. So now the problem becomes is it right that I have to pre-install the packages on nodes manually? And if that's true, what's the use case for --jars and --packages described in question 1? If that's wrong, how to use --jars and --packages to install the packages?