What standard are these backslash escape sequences using?

Viewed 49

I'm doing some text processing on Websters Unabridged Dictionary and have come across some escape sequences that don't conform to any standard (i.e., they are not HTML, CSS, Unicode, etc escape sequences) that I know of.

Some sample text:

<h1>Galore</h1>
<Xpage=610>

<hw>Ga*lore"</hw> <tt>(?)</tt>, <tt>n. & a.</tt> <ety>[Scot. <ets>gelore</ets>, <ets>gilore</ets>, <ets>galore</ets>, fr. <ets>Gael</ets>. <ets>gu le\'95r</ets>, enough; <ets>gu-</ets> to, also an adverbial prefix + <ets>le\'95r</ets>, <ets>le\'95ir</ets>, enough; or fr. Ir. <ets>goleor</ets>, the same word.]</ety> <def>Plenty; abundance; in abundance.</def>

They are all of the form \'xy where x, y are any digit in [0-9] or letter in [a-f]. Clearly, they agree in form with RTF escape sequences. However, the characters they are supposed to represent are nowhere near correct.

For the ones that appear in the data I want I have figured out that:

{
   "\'80": "Ç",
   "\'81": "ü",
   "\'82": "é",
   "\'83": "â",
   "\'84": "ä",
   "\'85": "à",
   "\'86": "å",
   "\'87": "ç",
   "\'88": "ê",
   "\'89": "ë",
   "\'90": "É",
   "\'91": "æ",
   "\'92": "Æ",
   "\'93": "ô",
   "\'94": "ö",
   "\'95": "ò",
   "\'96": "û",
   "\'97": "ù"
}

At first I thought that maybe it was a simple wrap-around error (each hexvalue xy is off by the same amount), but this is not true by looking at ç and ö and comparing their offset to the correct values or by noting that if Ç is \'80, then ü should be \'b5.

For completeness, all the values I found with the regex r"\\\'[\d\w]{2,2}" (74 in total) are:

\'3c
\'3e
\'80
\'81
\'82
\'83
\'84
\'85
\'86
\'87
\'88
\'89
\'8a
\'8b
\'8c
\'8d
\'90
\'91
\'92
\'93
\'94
\'95
\'96
\'97
\'9a
\'9c
\'a0
\'a1
\'a2
\'a3
\'a4
\'a6
\'a7
\'ab
\'ac
\'b5
\'b6
\'b7
\'b8
\'bd
\'be
\'bf
\'c3
\'c5
\'c6
\'c7
\'c8
\'c9
\'cb
\'cc
\'ce
\'cf
\'d0
\'d1
\'d2
\'d3
\'d4
\'d6
\'dc
\'dd
\'de
\'df
\'dh
\'eb
\'ed
\'ee
\'ef
\'f0
\'f4
\'f5
\'f6
\'f7
\'f8
\'fb

Can anyone tell me what standard these escape sequences are obeying? A link to a table or a library that will convert them into Unicode would be appreciated.


EDIT

Further processing revealed that:

{
   "\'d1": "Œ",
   "\'d2": "œ",
   "\'ee": "ã"
}

Unfortunately, it seems that while characters in \'80 - \'a5 conform to the IBM codepage 437, whoever made the document had decided to use a custom mapping for characters not in the original encoding, alas.

0 Answers
Related