reading Excel file getting unicodes

Viewed 353

I am reading an excel file with pandas.

When i open the file in microsoft excel then I got the output like this

enter image description here

when I see this file in libre office i got the output like this,

enter image description here

So while reading the excel file, I do the following code but i am not able to get rid of x000d

df = pd.read_excel('file.xlsx')
df = df.replace(r'\n',' ', regex=True)
df = df.replace(r'[^\x00-\x7F]+',  '', regex=True)

Also there can be more unicodes like this in whole file. The above code replaces all new lines from each cell.

1 Answers

Currently I am only able to solve this issue by finding these types of unicodes and then replacing them.

df.replace({r"_x([0-9a-fA-F]{4})_": ""}, regex=True)

Let me know if someone having better idea.

Related