How to decode hex encoded Cyrillic string?

Viewed 137

I have a hex encoded Cyrillic string "041E043F043B04300442".

How can I convert into a text string?

I tried this way:

codecs.decode('041E043F043B04300442', 'hex').decode('utf-16')
'Ḅ㼄㬄〄䈄'

But I'm getting wrong symbols. As I see from the Unicode symbols list, the first symbol should be a Cyrillic symbol:

U+041E  Cyrillic Capital Letter O

What am I doing wrong?

2 Answers

I had to use another codec:

codecs.decode('041E043F043B04300442', 'hex').decode('utf-16be')

Now it is being decoded fine.

utf-16 defaults to the machine's endian-ness unless a byte order mark (BOM, U+FEFF) is present. Your machine appears to be little-endian, but the data is big-endian:

>>> bytes.fromhex("041E043F043B04300442").decode('utf-16')
'Ḅ㼄㬄〄䈄'
>>> bytes.fromhex("041E043F043B04300442").decode('utf-16le')
'Ḅ㼄㬄〄䈄'
>>> bytes.fromhex("041E043F043B04300442").decode('utf-16be')
'Оплат'

(English) payment

With a correct BOM added, utf-16 can work:

>>> bytes.fromhex("FEFF041E043F043B04300442").decode('utf-16')
'Оплат'
Related