Replace multiple asterisks (***) DataFrame Python

Viewed 44

I need to replace in a dataframe the text:
J***ge by Jorge, I have tried many solutions but I always get this error

/usr/lib/python3.7/sre_parse.py in _parse(source, state, verbose, nested, first) 646 if item[0][0] in _REPEATCODES: 647 raise source.error("multiple repeat", --> 648 source.tell() - here + len(this)) 649 if item[0][0] is SUBPATTERN: 650 group, add_flags, del_flags, p = item0

error: multiple repeat at position 2

This is a summary of the DataFrame
It's like 5000 records

import pandas as pd

colors = {'first_name':  ['J***ge','Luis','Peter','Doug'],
          'second_name': ['Chavez','Ma**ani','Jhons','Leake']
         }

df = pd.DataFrame(colors, columns= ['first_name','second_name'])

print (df)

DataFrame Sample

3 Answers

Clean the data before importing to pandas data frame. Use dictionary comprehension to modify the dictionary, and use list comprehension to modify the list inside each dictionary item. Use re.sub to replace the asterisks.

Note that * has a special meaning in regular expressions. It is a modifier that says: repeat the previous character 0 or more times. Therefore, it needs to be escaped like so: \*, or put inside a character class like so: [*], which is essentially a class that consists of a single element - the asterisk. And [*]+ means an asterisk repeated 1 or more times.

import re
colors = {'first_name':  ['J***ge','Luis','Peter','Doug'],
          'second_name': ['Chavez','Ma**ani','Jhons','Leake']
         }
colors = {k: [re.sub(r'[*]+', '', s) for s in lst] for k, lst in colors.items()}
print(colors)
# {'first_name': ['Jge', 'Luis', 'Peter', 'Doug'], 'second_name': ['Chavez', 'Maani', 'Jhons', 'Leake']}

Thank you Timur Edit your coe is:

import re
colors = {'first_name':  ['J***ge','Luis','Peter','Doug'],
          'second_name': ['Chavez','Ma**ani','Jhons','Leake']
         }
colors = {k: [re.sub(r'[***]+', 'or', s) for s in lst] for k, lst in colors.items()}

colors = pd.DataFrame(colors, columns = ['first_name','second_name'])
colors

But replace all * Solve one datum, but not the others

The Name Ma**ani replace to Maorani, the correct is Mamani

I already found the way to do it. First transform the data of the DataFrame to a list and then change them. If someone finds a way with fewer lines, please let me know.

#The source DataFrame
import pandas as pd

colors = {'first_name':  ['J***ge','Luis','Peter','Doug'],
          'second_name': ['Chavez','Ma**ani','Jhons','Leake']
         }

df = pd.DataFrame(colors, columns= ['first_name','second_name'])

print (df)
print(df.columns)

# Replacing the manifolds ***
# clean the database. // Select the column
df1 = df[['first_name']]

df_up = df1['first_name'].tolist() #convert column to list
s = str(df_up) # the list to string
new_s = s.replace('J***ge','Jorge') # replacing
output = eval(new_s) # replacing
output = pd.DataFrame(output, columns = ['first_name_rem']) # the list replaced to Dataframe
output = pd.concat([df, output], axis=1) #Concatenate the DataFrame
print(output)
Related