Let's say I want to remove the word "tree" in every string in a Pandas dataframe column.
I would specify the substring(s) I want removed in a list. And then use replace and join on the column, as per below:
remove_list = ['\tree\s']
df['column'] = df['column'].str.replace('|'.join(remove_list ), '', regex=True).str.strip()
The reason I add a \s to tree is because there may be words like treehouse or backstreet. So I want to replace the word only if it ends with a space, so that I don't end up with words like "house" or "backst".
However I noticed that when I run this code, it misses "tree"s that are at the end of the string, because there is no space after it. Hence, it doesn't get removed. Any idea on how I can account for those?