I'm learning to use R, so please bear with me.
I have a dataset of google play store apps (master_tib). Each row is a play store app. There's a column titled description that contains text on what the app does.
master_tib
App Description
App1 Reduce your depression and anxiety
App2 Help your depression
App3 This app helps with Anxiety
App4 Dog walker app 3000
I also have a df of tags (master_tags) that contains words of importance I've predefined. There is a single column titled tag and each row contains a single tag.
master_tag
Tag
Depression
Anxiety
Stress
Mood
My goal is to tag apps from the master_tib df with the tags in master_tags df based on the presence of the tag in the description. It will then print the tags in a new column. The final result would be a master_tib df that looks like this:
App Description Tag
App1 Reduce your depression and anxiety depression, anxiety
App2 Help your depression depression
App3 This app helps with anxiety anxiety
App4 Dog walker app 3000 FALSE
Below is what I've done so far using a combination of str_detect and mapply:
# define function to use in mapply
detect_tag <- function(description, tag){
if(str_detect(description, tag, FALSE)) {
return (tag)
} else {
return (FALSE)
}
}
index <- mapply(FUN = detect_tag, description = master_tib$description, master_tags$tag)
master_tib[index,]
Unfortunately, only the first tag is being passed through.
App Description Tag
App1 Reduce your depression and anxiety depression
Instead of the desired:
App Description Tag
App1 Reduce your depression and anxiety depression, anxiety
I haven't gotten as far as printing the results into a new column. Would love to hear anyone's insight or thoughts and apologize in advance for my poor R skills.