I have the following DataFrame:
df = pd.DataFrame({'phrase':['websocket internet is loading','foo bar kangeroo bunny','websocket funny internet','scrape the internet with websocket','another one']})
df
phrase
0 websocket internet is loading
1 foo bar kangeroo bunny
2 websocket funny internet
3 scrape the internet with websocket
4 another one
I am trying to use regex with Pandas' str.contains() to match phrases with multiple words, but not requiring the exact sequence of those words to match. I would like to match using a list of strings:
['websocket internet', 'foo bunny'].
Expected output:
0 True
1 True
2 True
3 True
4 False
I know I can implement regex like so:
df['phrase'].str.contains(r'^(?=.*\bfoo\b)(?=.*\bbunny\b).*$')
df['phrase'].str.contains(r'^(?=.*\bwebsocket\b)(?=.*\binternet\b).*$')
But what if I have a large list of match terms? How would I format all of those strings to the necessary regex, and is there a way for me to implement multiple regex in one str.contains() function?