I am working on an older project (compiled with UNICODE defined) and came across a problem within the .rc. For example, a static text element which includes “©” defined in a DIALOGEX resource by
LTEXT "Copyright ©”,IDC_COPYRIGHT_STATIC,7,154,110,8
The resource file, probably created by MSVC application wizard many years ago and migrated forward with each release, now looks like this:
#if !defined(AFX_RESOURCE_DLL) || defined(AFX_TARG_ENU)
LANGUAGE LANG_ENGLISH, SUBLANG_ENGLISH_US
#pragma code_page(1252) //present for over 10 years
#endif
For many years the © display correctly but recently appeared as “Å©” or even “½¿”. Obviously, an encoding issue, but I needed to understand how and why before making changes. So, after researching, these three properties in the .rc play a part in the bug and the encoding:
- The presence or absence of “#pragma code_page(…)” in the .rc
- The encoding used to Save with Encoding… the .rc file
- Save with Encoding… .rc “with signature” or “without signature” (meaning BOM?)
As an empirical test, changing these things in the .rc and looking at the result text in dialogue
| #pragma code_page(…) | Save with Encoding | Signature(BOM) | Text in Dlg |
|---|---|---|---|
| code_page(1252) | Original file | n/a | Å© |
| code_page(1252) | Windows 1252 | n/a | © |
| code_page(1252) | UTF-8 65001 | BOM | Å© |
| code_page(1252) | UTF-8 65001 | No BOM | Å© |
| code_page(65001) | Windows 1252 | n/a | © |
| code_page(65001) | UTF-8 65001 | BOM | © |
| code_page(65001) | UTF-8 65001 | No BOM | © |
| No code_page in .rc | UTF-8 65001 | BOM | © |
| No code_page in .rc | UTF-8 65001 | No BOM | Å© |
I can explicitly Save with Encoding all .rc files encoding as Windows (1252) OR encoding as UNICODE UTF-8 with signatures (and delete the #pragma code_pages). The specific bug will go away, but is this the best solution?
It seems switching from Windows 1252 to UNICODE UTF-8 is a step forward and the right way to go long term. Is there any problem with this? Or better solutions?
