how generate random values for groups by arithmetic condition in R

Viewed 189

I have dataset with such structure

mydata=structure(list(supps = c("KR", "KR", "KR", "KR", "KR", "KR", 
"KR", "KR", "KR", "KR", "aeroclub", "aeroclub", "aeroclub", "aeroclub", 
"aeroclub", "aeroclub", "aeroclub", "aeroclub", "aeroclub", "aeroclub"
), date = c("01.05.2021", "01.06.2021", "02.05.2021", "02.06.2021", 
"03.05.2021", "03.06.2021", "04.05.2021", "04.06.2021", "05.05.2021", 
"05.06.2021", "01.05.2021", "01.06.2021", "02.05.2021", "02.06.2021", 
"03.05.2021", "03.06.2021", "04.05.2021", "04.06.2021", "05.05.2021", 
"05.06.2021"), turnover = c(0, 0, 32159.00888, 25220.0027, 0, 
0, 245312.682, 189901.1224, 0, 0, 1531959.833, 1591612, 1834696.667, 
1885169, 1871615.167, 1823398, 4891342, 5253701.167, 0, 0), fee = c(0, 
0, 651, 37, 0, 0, 2341, 7548, 0, 0, 40519.5, 30415, 34767.66667, 
39289, 39175.66667, 45798, 94819.5, 116803.1667, 0, 0), comiss = c(0, 
0, 764.81, 537.67, 0, 0, 8578.25, 6198.115, 0, 0, -2023.41, -1941.67, 
-550.82, 1323.23, -1029.47, -638.47, -1034.58, -1332.95, 0, 0
), intencive = c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 12, 26.4, 1945.8, 
2199.48, 3740.76, 6499.2, 32188.68, 42337.44, 0, 0)), class = "data.frame", row.names = c(NA, 
-20L))

I need for each group by supps column(KR and aeroclub) for vriables turnover fee comiss intencive calculate value by next condition. For example we take KR and turnover variable. The last 2 values belong dates 03.06.2021-04.06.2021. if the most recent value is greater than the previous one, then calculate sum of value 189901+0=189901. Then for each variable for dates 05.06.2021-08.06.2021(4 days) generate random value. This calculated Sum 189901+(2%-10%) from it, in random order. To be more clear For exampe output (for turnover variable)

05.06.2021  189901+2%=193699,02
06.06.2021 189901+10%=208891,1
07.06.2021  189901+6%=208891,1
08.06.2021 189901+7%=203194

But sometimes can be that last values is negative. for example. group=aeroclub. Variable comiss, last 2 values 03.06.2021-04.06.2021 at 04.06.2021 value -1332, but at 03.06.2021 value -632, so at 04.06.2021 value less then at 03.06.2021. We sum these value -1332+-632=-1954 but then we didn't plus to sum, we subtrack -1954-(2%-10%)in random order. So for this group by comiss desired output

05.06.2021  -1954-2%=-1914,92
06.06.2021 -1954-7%=-1817,22
07.06.2021  -1954-6%=-1836,76
08.06.2021 -1954-8%=-1797,68

How can i do it correct?

1 Answers

The following answer assumes few things which weren't entirely clear from the question:

  1. The calculations are done starting from column no.3 to the last column
  2. When there is 0, it is kept as 0. No random % is added. Though you may change that if you want.
  3. There are few instances when the two consecutive values have different signs. For those cases, the rule for the latest value is applied following the question.
#storing the unique category of supps
col_supps <- unique(mydata$supps)
#storing the columns for which the calculations will be done
col_names <- colnames(mydata)[3:ncol(mydata)]

#the data frame which will contain the output
output_df <- data.frame()
#Iterating over different supps values

for (x in col_supps) {
#storing one type of supps in a temporary data frame
  mydata %>%
    filter(supps %in% x)-> temp
  temp1<- temp

#temp will act as a reference frame, in temp1 values will be updated
  
#Now, iterating over columns which we need
  
  for (y in col_names) {
    i<- 1
#with while loop, we will iterate over each elememnt of the column and save the result in temp1
    while (i<=(nrow(temp)-1)) {
      if(temp[i+1,y]>0 & temp[i+1,y]>=temp[i,y]){
        temp1[i+1,y] <- (temp[i,y]+temp[i+1,y]) * (runif(1,1.02,1.1))
      }else if(temp[i+1,y]<0 & temp[i+1,y]<=temp[i,y]){
        temp1[i+1,y] <- (temp[i,y]+temp[i+1,y]) * (runif(1,1.02,1.1))
      }else if(temp[i+1,y]>0 & temp[i+1,y]<temp[i,y]){
        temp1[i+1,y] <- (temp[i+1,y]) * (runif(1,1.02,1.1))
      }else if(temp[i+1,y]<0 & temp[i+1,y]>temp[i,y]){
        temp1[i+1,y] <- (temp[i+1,y]) * (runif(1,1.02,1.1))
      }else if(temp[i+1,y]==0){
        temp[i+1,y] <- 0
      }
      i <- i+1
    }
  }
#saving the output in the output data frame before repeating the process for another type of supps
  output_df %>%
    bind_rows(temp1) -> output_df
  
}

Now output_df will have the final output which you desire. If you want reproducibility in the random value, you can set.seed() under the while loop. If that is not wanted, then you can proceed as it is.

Related