I have a DataFrame where some values are stored as "Miami–Fort Lauderdale" and "Minneapolis–Saint Paul" with longer hyphen "–" (not short dash "-"). I am trying to remove them with regex in Windows command prompt, but it's not working properly.
- directly typing the hyphen as below does not work (werid enough):
XXX.replace(to_replace=r'\–', value=' ', regex=True)
XXX.replace(to_replace='–', value=' ')
and gives unchanged "Miami–Fort Lauderdale" and "Minneapolis–Saint Paul". Thus, I suppose for some reason cmd does not recognize hyphen.
- the general form is "lowercase letter + hyphen + uppercase letter" so I also tried
XXX.replace(to_replace=r'(?=[a-z]+)\W(?=[A-Z]+)', value=' ', regex=True)
interestingly this gives unchanged "Miami–Fort Lauderdale" and "Minneapolis–Saint Paul"
- however, the following will work
XXX.replace(to_replace=r'\W(?=[A-Z]+)', value=' ', regex=True)
and gives desired "Miami Fort Lauderdale" and "Minneapolis Saint Paul". But the problem is that this messes up other values like "Washington, D.C." into "Washington, D C." (apparently).
=====================================================
I eventually solved this by
XXX.replace(to_replace=r'\W(?=\w+\s)', value=' ', regex=True)
but I still wonder how Regex recognizes the letter before hyphen "–". It appears to me as if for some reason, a letter right before hyphen is not considered as a letter?