Using an extended-ASCII collation table in C for Win32

Viewed 135

In an existing app an API was used to convert all ASCII characters to an uppercased byte, for sorting purposes. cafe == CAFE == Café == CAFÉ. The characters e, é and E all became the character E, in a sorting name. So the value of table[137], representing the byte value of an é, was 69 ("E").

I've performed a few tests with an older, required Win32 API, which converted a whole string but didn't convert the character é to the character É nor to the expected character E.

How can this be done with an older Windows API?

1 Answers

Solved, by finally almost doing what paulsm4 suggested.

The input codepage is, or should be, known: 850. 858 may be supported too, but let's assume 850. I already had to convert a few known characters to HTML, so I know which characters are occuring in the input text.

Instead of a table I'm using a function for the (14) exceptions (and a-z):

case 'ë':
case 'é':
   return 'E';
case 'ï':
   return 'I';

The original dynamic API (different operating system) or toupper() may be best indeed, albeit one may argue that é should become E or É.

With a limited and known number of exceptions a function, using static values, is an option too.

Finally one could use the other, original operating system to produce code for a static, full table with upto 256 elements. So you don't have to type 512 numbers manually and correctly.

The other, rare, old operating system has an API which returns <= 256 values, using the current codepage:

if (name[i]<=length_table)
   sortname[i]=table[i];
else
   sortname[i]=name[i];

New code:

sortname[i]=table(i);

Speed is not an issue, but the new function first takes care of more usual characters. It's customized, so I won't poost it here. In theory toupper() could be used if the character is <= 127 there.

Related