I have a number of strings (CIGARs) that I am trying to sum the numbers that occur before the number preceding "I". The position that "I" occurs is highly variable but always has a number before it.
Here is a sample df:
df <- data.frame(String = c("220M1I","10I200M","5M2D1I20M","22M5D2M3I5M"))
My desired output looks like:
String Sum_prior
1 220M1I 220
2 10I200M 0
3 5M2D1I20M 7
4 22M5D2M3I5M 29
I have a partial solution which can't handle >1 digit numbers prior to "I" which is problematic.
sum_fun <- function(x) {
str_match_all(x, "\\d+(?!I)") %>%
unlist() %>%
as.numeric() %>%
sum()
}
then applying to df:
df <- df %>% rowwise() %>% mutate(output = sum_fun(String))
df
String output
<chr> <dbl>
1 220M1I 220 #Good
2 10I200M 201 #The 1 in 10 is being included
3 5M2D1I20M 27 #Don't want last 20 included
4 22M5D2M3I5M 34 #Don't want last 5 included
But I can't figure out how to adapt the regex to ignore all numbers immeadiately prior to "I" and sum all other numbers before "I".
A more advanced example I need (but less important), is to calculate the cumulative number when there is more than one "I" - the first occurrence is as above (output_1), but the second (or more) (output_2) example includes the preceeding "I" number.
df2 <- data.frame(String =c("5M10I200M20I","100M2D3I105M1I10M")
String Output_1 Output_2
1 5M10I200M20I 5 215
2 100M2D3I105M1I10M 102 210
Any help is appreciated.