I have a data.frame that is >250,000 columns and 200 rows, so around 50 million individual values. I am trying to get a breakdown of the variance of the columns in order to select the columns with the most variance.
I am using dplyr as follows:
df %>% summarise_if(is.numeric, var)
It has been running on my imac with 16gb of RAM for about 8 hours now.
Is there a way top allocate more resources to the call, or a more efficient way to summarise the variance across columns?