Conditional row count based on two variables (R)

Viewed 68

I have a problem I can't wrap my head around at the moment. I have a large dataframe that looks something like this:

df <- data.frame( 
         Marker = c("", "", "", "start_tone", "", "", "start_trial", "", "", "", "", "", "", "", "start_tone", "", "", "", "start_trial", "", "", "", "", "", ""), 
         size=c(3, NA, -1, -1, 4, -1, -1 , 3.5, -1, -1, 4, -1, -1, NA, 4, -1, -1, 2, -1, -1, -1, -1, -1, 4.5, -1))

df

What I want is to count the number of samples starting from the "start_trial" marker. I want that count to go both directions, in negative and positive values. Furthermore, I want that negative count to stop an n amount of samples (say 2) before the "start_tone" marker. In the positive direction, I want that count to stop n (here 2) samples before the next "start_tone" marker (where the negative count starts). Lastly, I want a new column that keeps track of the trials around each "start_trial" marker. So, it should look something like the following:

df2 <- data.frame(SampleCount = c("", -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, -6, -5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5, 6), 
         Marker = c("", "", "", "start_tone", "", "", "start_trial", "", "", "", "", "", "", "", "start_tone", "", "", "", "start_trial", "", "", "", "", "", ""), 
         size=c(3, NA, -1, -1, 4, -1, -1 , 3.5, -1, -1, 4, -1, -1, NA, 4, -1, -1, 2, -1, -1, -1, -1, -1, 4.5, -1),
         trial=c("", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial 1", "trial1", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2", "trial 2" ))

df2

Thanks a lot for helping out!

1 Answers

Use cumsum and lead to create the groups, and then use row_number to create the sample count:

df %>% 
  group_by(trial = cumsum(lead(Marker, 2, default = "") == "start_tone")) %>% 
  mutate(n = ifelse(any(Marker == "start_trial"), row_number()[Marker == "start_trial"], NA),
         sampleCount = row_number() - n) 

output

   Marker         size trial     n sampleCount
   <chr>         <dbl> <int> <int>       <int>
 1 ""              3       0    NA          NA
 2 ""             NA       1     6          -5
 3 ""             -1       1     6          -4
 4 "start_tone"   -1       1     6          -3
 5 ""              4       1     6          -2
 6 ""             -1       1     6          -1
 7 "start_trial"  -1       1     6           0
 8 ""              3.5     1     6           1
 9 ""             -1       1     6           2
10 ""             -1       1     6           3
11 ""              4       1     6           4
12 ""             -1       1     6           5
13 ""             -1       2     7          -6
14 ""             NA       2     7          -5
15 "start_tone"    4       2     7          -4
16 ""             -1       2     7          -3
17 ""             -1       2     7          -2
18 ""              2       2     7          -1
19 "start_trial"  -1       2     7           0
20 ""             -1       2     7           1
21 ""             -1       2     7           2
22 ""             -1       2     7           3
23 ""             -1       2     7           4
24 ""              4.5     2     7           5
25 ""             -1       2     7           6
Related