I have arabic text in a dataframe and I want to remove the letter و from all words that start with this letter. I tried to do this:
def clean(text_string):
space_pattern = '\bو'
parsed_text = re.sub(space_pattern, '', text_string)
return parsed_text
and then:
df['tidy_tweet'] = np.vectorize(clean)(df['tidy_tweet'])
but when I run it, nothing changes. It's as if I didn't do anything at all!
Example:
Input: هيه الهزه الحقيقيه وتخافون الهزه وماتخافون الهزه اعملها نظامكم الهمجي
Desired output: هيه الهزه الحقيقيه تخافون الهزه ماتخافون الهزه اعملها نظامكم الهمجي