I have a set of strings that I want to check if they contain any character that is dynamically created in a separate data frame. For example:
library(dplyr)
library(stringr)
#Create data frame of Values to check against
df <- data.frame(Value = 33:53) #these values will be dynamic
#Add in Utf8 conversion
df <- df %>% rowwise() %>% mutate(Utf8 = intToUtf8(Value))
#create data frame of strings to check
df_check <- data.frame(string = c("CDEFFFFFFFEFADFFFFF","CDEFFFFFFF&FADFFFFF"))
I now want to check if the strings contain any of the characters that are in df, and output either TRUE or FALSE in a new column. I don't want to use a standard regular expression to check if any of the characters are there, because the length of df can be variable and certain characters can be omitted.
I can check individual characters like so (which gives me the desired output but for only one character):
df_check <- df_check %>% rowwise() %>% mutate(contains_chr = list(str_detect(string,"&")))
> string contains_chr
> CDEFFFFFFFE!FADFFFFF FALSE
> CDEFFFFFFF&!FADFFFFF TRUE
How can I make the 'contains_chr' column TRUE or FALSE based on whether any of the characters in df appeared? I tried this but it fails:
df_check <- df_check %>% rowwise() %>% mutate(contains_chr = list(str_detect(string,df)))
Error: Problem with `mutate()` column `contains_chr`.
ℹ `contains_chr = list(str_detect(string, df))`.
x no applicable method for 'type' applied to an object of class "c('rowwise_df', 'tbl_df', 'tbl', 'data.frame')"
ℹ The error occurred in row 1.
I don't know how to check for each of the characters in df$Utf8. I probably could put it into a for-loop for each row in df, but there are millions of strings and this seems like it will be quite slow.