Encoding issue during reading excel file in Python

Viewed 291

I use read_excel from pandas library to read excel content and convert it to JSON. I am struggling with encoding issue. Non english characters are encoded like "u652f\u63f4\u8cc7\u8a0a". How can I resolve this issue? I tried

wb = xlrd.open_workbook(excel_filePath, encoding_override='ISO-8859-1')
new_data = pd.read_excel(wb)

Also

with open(excel_filePath, mode="r", encoding="utf-8") as file:
  new_data = pd.read_excel(excel_filePath)

I tried this code with encodings like: utf-8, utf-16, utf-16, latin1...

1 Answers

From the docs of the json module:

The RFC requires that JSON be represented using either UTF-8, UTF-16, or UTF-32, with UTF-8 being the recommended default for maximum interoperability.

As permitted, though not required, by the RFC, this module’s serializer sets ensure_ascii=True by default, thus escaping the output so that the resulting strings only contain ASCII characters.

Maybe surprising that in this day-and-age the module defaults to escaping non-ASCII (probably for backwards compatibility), so just override that behavior with ensure_ascii=false:

with open(json_filePath, 'w') as f:
    json.dump(new_json, f, ensure_ascii=False)
Related