I have a dataset like:
ID val1 val2 val3 val4
1 4 9 10 16
2 1.5 6 2.3 99
3 8 7 7 10
I would like to check whether the number of columns (i.e., val columns) is less than 6 and if that is the case, I want to randomly select the number of columns left from the existing columns and add them again to the dataset.
In the case above, the number of columns left is 2 (6 - 4 columns of val). In this case, I would like to select 2 random columns from the val columns and add them to the dataset. one possible solution would be:
ID val1 val2 val3 val4 val2 val1
1 4 9 10 16 9 4
2 1.5 6 2.3 99 6 1.5
3 8 7 7 10 7 8
Columns val2 and val1 are randomly selected and added to the dataset.
The problem that I'm facing is how to select random columns. I know how to select random rows by using sample_n function, but I couldn't find any function to select random columns.
What I did so far is:
t <- read.csv("path", header=TRUE) # load file
numCols <- 6
cc <- ncol(t[,-1]) #no need for ID column
if(cc < numCols){
# I need some function to select random columns
}