I am working with the R programming language. Suppose I have the following dataset:
var_1 <- c("A","B")
var_1 <- sample(var_1, 1000, replace=TRUE, prob=c(0.3, 0.7))
var_1 <- as.factor(var_1)
var_2 <- c("AA","BB", "CC")
var_2 <- sample(var_2, 1000, replace=TRUE, prob=c(0.2, 0.1, 0.7))
var_2 <- as.factor(var_2)
var_3 <- c("AA1","BB1")
var_3 <- sample(var_3, 1000, replace=TRUE, prob=c(0.5, 0.5))
var_3 <- as.factor(var_3)
my_data = data.frame(var_1, var_2, var_3)
my_data$var4 = rnorm(1000,10,10)
my_data$var5 = rnorm(1000,10,10)
my_data$var6 = rnorm(1000,10,10)
head(my_data)
var_1 var_2 var_3 var4 var5 var6
1 B AA AA1 6.960184 17.191858 -11.977489
2 B CC BB1 14.672173 9.953185 4.712377
3 B CC BB1 -5.211582 3.930513 25.637752
4 B CC AA1 14.252484 2.898963 7.629264
5 A CC AA1 16.387029 15.608298 18.234116
6 A CC BB1 7.433072 28.338435 -1.726043
My Question: In the above data, there are 3 factor variables with a total of 12 possible combinations of these factors. I am trying to extract each of these combinations and create 12 new datasets, dedicated to each one of these combinations.
I tried the following code in R:
lst1 <- split(my_data, my_data[c("var_1", "var_2", "var_3")], drop = TRUE)
I then tried to extract each of these dataset form this list:
list2env(lst1,envir=.GlobalEnv)
Can someone please tell me if what I have done is correct? If there were more variables and more factor combinations - is there a more efficient way to solve this problem?
Thanks!