I have a string column of narratives. Each narrative is basically an essay. I want to take a subset of the df where certain phrases exist. The current method isn't working as intended. I'm filtering rows that don't contain the phrase exactly or just contains a subset of the phrase.
I've tried the following:
phrase = ['went to the store to buy an apple', 'corner of the street', 'fbi most wanted']
df['text'].str.contains(r'\b{}\b'.format('|'.join(phrase)), re.IGNORECASE, regex=True)
Not including an example because really just looking for a code review more than anything. The method above should look through the column text to see if those phrases exist, correct? Or am I missing something?