Resampling cross-sectional time series data in R

Viewed 195

I'm dealing with cross-sectional time series data (many DIFFERENT individuals over time). At the individual level, each person has a quantity of a good demanded. This data is unbalanced with respect to how many individuals are in each period. For each time period, I've aggregated the individual data into a single time series. Example data structure below

Cross-Section Time Series

Time | Person | Quantity
----------------------
11/18| Bob    | 2
11/18| Sally  | 1    
11/18| Jake   | 5
12/18| Jim    | 2   
12/18| Roger  | 8

Time Series

Time | Total Q
-------------
11/18| 8      
12/18| 10    

What I want to do for each period is resample (with replacement) the individual quantity, aggregate across the individuals, iterate X amount of times, and then get an mean and standard error from the bootstrap.

The end result should look like

Time | Total Q | Boot Strap Total Mean  
-------------------------------------
11/18| 8       | 8.5 
12/18| 10      | 10.05 

Here is some code to create example sample data:

library(tidyverse)

set.seed(1234)

Cross_Time = data.frame(x) %>%
     mutate(Period = sample(1:10, 50, replace=T),
            Q=rnorm(50,10,1)) %>%
     arrange(Period)

Timeseries = Cross_Time %>%
group_by(Period) %>%
summarize(Total=sum(Q))

I know this is possible in R, but I'm at a loss as to how to code it or what the right questions I need to ask are. All help is appreciated!

1 Answers

We may do the following:

X <- 1000
Cross_Time %>% group_by(Period) %>%
  do({QS <- colSums(replicate(sample(.$Q, replace = TRUE), n = X))
  data.frame(Period = .$Period[1], `Total Q` = sum(.$Q), Mean = mean(QS), `Standard Error` = sd(QS))})
# A tibble: 10 x 4
# Groups:   Period [10]
#    Period Total.Q  Mean Standard.Error
#     <int>   <dbl> <dbl>          <dbl>
#  1      1    28.8  28.8          0.284
#  2      2    35.9  35.8          0.874
#  3      3   109.  109.           3.90 
#  4      4    48.9  48.9          2.16 
#  5      5    20.2  20.2          0.658
#  6      6    59.0  58.8          3.57 
#  7      7    88.7  88.6          2.64 
#  8      8    22.7  22.7          1.04 
#  9      9    47.7  47.7          2.46 
# 10     10    27.9  27.9          0.575

I think the code is quite self-explanatory. In every group we resample it's values with replacement X times with replicate and compute the two desired statistics. It's also straightforward to add any others!

Related