Say I have a vector A (or a variable A in a dataframe say df) with following values
A <- c(90L, 100L, 5L, 15L, 16L, 2L, 20L, 25L, 2L, 40L, 16L, 16L, 32L, 51L, 52L)
A
#> [1] 90 100 5 15 16 2 20 25 2 40 16 16 32 51 52
df <- data.frame(A = A)
Created on 2021-05-11 by the reprex package (v2.0.0)
Now I want to divide these values into two States say 0 and 1 based on the following criteria
- By default first value will be
0state. - If the value has dropped by more than
80%of its previous value (e.g. in third row from 100 to 5 i.e. 95% drop), its state changes to1from0. - Here the difficult part arrives. All values following this dropped value will remain in
1state until it rises 50% of that value (100) i.e. 50 again. It rises above 50 in 14th row. So value in 14th row will be state0again - Since last value is not a decrease by more than 80% of previous value i.e. 51 it will be same state i.e.
0.
So basically I am trying to get an output like this
A State
1 90 0
2 100 0
3 5 1
4 15 1
5 16 1
6 2 1
7 20 1
8 25 1
9 2 1
10 40 1
11 16 1
12 16 1
13 32 1
14 51 0
15 52 0
BaseR or tidyverse approach will be doing fine for me.
Where I am stuck is actually retrieving the threshold value (100 or 50%) till 14th row as you can see it further drops by 80% again two times, once in row 6th and again row 9th.
One more test case can be
A State
1 90 0
2 100 0
3 5 1
4 15 1
5 16 1
6 2 1
7 20 1
8 25 1
9 2 1
10 40 1
11 16 1
12 16 1
13 32 1
14 51 0
15 52 0
16 60 0
17 10 1
18 20 1
19 5 1
20 30 1
21 31 0
22 50 0
23 100 0
Explanation
benchmarking both answers on df <- data.frame(A = sample(1:100, 100000, T))
Unit: microseconds
expr min lq mean median uq max neval
BlueVoxe() 214547.5 271455.60 298551.084 300763.90 309823.3 499692.5 100
Ben() 4.2 4.85 7.115 5.35 9.8 11.3 100
