While I had received some great feedback on my previous post, I believe that my original question was not entirely clear and hence the answers did not generate the desired outcome.
I have a long vector of a character variable strings with about 600K observations having 800 unique string values. I am trying to narrow down these 800 unique strings to about 20 unique strings based on another vector of important string variables.
Here is an example:
col1 <- c("CORE_I5-xxxx_6C_VPRO", "A6-xxxx_MB", "CORE_I7-xxxx_4C_VPRO_MB", "INTEL_CORE_I3_MB", NA)
col2 <- c("CORE_I5_VPRO", NA, "CORE_I7_VPRO", "INTEL_CORE_I3", NA)
The new column (col2) has been created from the old column (col1) based on the following character variable (V) only by retaining the strings included in V:
V <- c("CORE", "INTEL", "I5", "I7", "I3", NA)
I have tried the following code but it is only giving me part of the strings, but not all the elements in each observation.
library(stringr)
col2 <- str_extract(col1, paste(V, collapse="|"))
I have also tried the suggestions to my previous post but unfortunately I am not getting the desired output. Thank you all for the help!