How can i speed up this loop in R?

Viewed 62
set.seed(155656494)

#setting parameter values
n<-500
sdu<-25
beta0<-40
beta1<-12

# Running the simulation again

# create the x variable outside the loop since it’s fixed in  
# repeated sampling 
x2 <- floor(runif(n,5,16))

# set the number of iterations for your simulation (how many values 
# of beta1 will be estimated)
nsim2 <- 10000000

# create a vector to store the estimated values of beta1 
vbeta2 <- numeric(nsim2) 

# create a loop that produces values of y, regresses y on x, and  
# stores the OLS estimate of beta1

for (i in 1:nsim2) {
  y2 <- beta0 + beta1*x2 + 0.2*x2 + rnorm(n,mean=0,sd=sdu)  
  model2 <- lm(y2 ~ x2)   
  vbeta2[i] <- coef(model2)[[2]] 
}

mean(vbeta2)

The above is a simple linear regression model that has 10 million iterations. I looking for help with speeding up the loop. This code basically runs as y2 <- beta0 + beta1x2 + 0.2x2 + rnorm(n,mean=0,sd=sdu), which will then be used to calculate the mean of vbeta2

1 Answers

You could use profvis to determine where is processor time spent:

library(profvis)

profvis({
  set.seed(155656494)
  
  #setting parameter values
  n<-500
  sdu<-25
  beta0<-40
  beta1<-12
  
  # Running the simulation again
  
  # create the x variable outside the loop since it’s fixed in  
  # repeated sampling 
  x2 <- floor(runif(n,5,16))
  
  # set the number of iterations for your simulation (how many values 
  # of beta1 will be estimated)
  nsim2 <- 1000
  
  # create a vector to store the estimated values of beta1 
  vbeta2 <- numeric(nsim2) 
  
  # create a loop that produces values of y, regresses y on x, and  
  # stores the OLS estimate of beta1
  
  for (i in 1:nsim2) {
    y2 <- beta0 + beta1*x2 + 0.2*x2 + rnorm(n,mean=0,sd=sdu)  
    model2 <- lm(y2 ~ x2)   
    vbeta2[i] <- coef(model2)[[2]] 
  }
  
  mean(vbeta2)
 
})

The result shows that most of the time is spent evaluating the linear regression model : enter image description here

As suggested by @maarvd, you could paralellize to speed this up. However parallelizing each single calculation won't be efficient because one calculation is too fast (~0.5 ms), so you'll have to distribute chuncks of many thousand calculations per worker. I agree with @Allan Cameron, is this worth the effort?

Related