I have thousands of data frames and I want to parallelise their analysis into slurm. Here I am providing a simplified example:
I have an Rscript that I call: test.R
test.R contains these commands:
library(tidyverse)
df1 <- tibble(col1=c(1,2,3),col2=c(4,5,6))
df2 <- tibble(col1=c(7,8,9),col2=c(10,11,12))
files <- list(df1,df2)
for(i in 1:length(files)){
df3 <- as.data.frame(files[1]) %>%
summarise(across(everything(), list(mean=mean,sd=sd)))
write.table(df3, paste0("df",i))
}
Created on 2022-04-15 by the reprex package (v2.0.1)
I want to parallelise the for loop and analyse each data frame as a separate job. Any help, guidance, and tutorials are appreciated.
Would the array command help?
#!/bin/bash
### Job name
#SBATCH --job-name=parallel
#SBATCH --ntasks=1
#SBATCH --nodes=1
#SBATCH --time=10:00:00
#SBATCH --mem=16G
#SBATCH --array=1-2
module load R/4.1.3
Rscript test.R $SLURM_ARRAY_TASK_ID