I have some old text documents that I'm trying to convert into UTF8 with Python. It has a .DOC extension, and it was written using a Tandy 1000. I haven't been able to find what kind of encoding it uses, so I'm thinking it was some Deskmate proprietary encoding. The majority of the text is ascii symbols, and a small minority are non-ascii characters. For example, see this xxd output:
2073 6f75 6c73 a9a9 a920 636f 6d70 6172 souls... compar
Using a printed version of the document, I can map the non-ascii bytes to their unicode characters. For example, in the above xxd output, 0xa9 is an em dash. There are many other examples of these non-ascii characters that I can manually map to utf8.
Any suggestions on the most elegant approach to decoding these characters? Right now I have the file loaded in a Python bytearray, and I'm wondering if there's a away to give the .decode("utf8") function a map to show that, say, 0xa9 is 0x2014 (i.e., em dash)?
For reference, here are the outputs of chardetect:
$ chardetect *DOC
file1.DOC : ascii with confidence 1.0
file2.DOC : ISO-8859-1 with confidence 0.7266894615606104
file3.DOC : ascii with confidence 1.0
file4.DOC : ascii with confidence 1.0
file5.DOC : Windows-1252 with confidence 0.7184626436781609
file6.DOC : Windows-1252 with confidence 0.7278633198811477
file7.DOC : Windows-1252 with confidence 0.726540670123209
file8.DOC : ISO-8859-1 with confidence 0.7252491632577166
file9.DOC : Windows-1252 with confidence 0.7268076883962168
file10.DOC : Windows-1252 with confidence 0.7252083189387135
file11.DOC : Windows-1252 with confidence 0.7270349029424267
file12.DOC : ISO-8859-1 with confidence 0.73
file13.DOC : KOI8-R with confidence 0.8191677051323929