I have a pyspark dataframe with names like:
- J.J. Scott
- J. S. Joyce
- RV. Bradley Carter
Some of them contain dots and spaces between initials and some do not. How can they be converted to:
- JJ Scott
- JS Joyce
- RV Bradley Carter
(with no dots and spaces between initials and 1 space between initials and name)
I tried using the following but it only replaces dots and doesn't remove spaces between initials:
names_modified = names.withColumn("name_clean", regexp_replace("name", r"\.",""))
Thanks!