As a foreword, I don't doubt that my knowledge of Unicode and character encoding is incomplete. Sorry for any misrepresentations.
I have recently been developing an application that will deal with incoming data in various encodings. These are, for example, UTF-8, ASCII, SHIFT-JIS, etc.
Now my understanding is that for a given encoding, a 'mapping exists' between it and a given Unicode value. Take for example the character 'あ'.
The Unicode value will be U+3042, the UTF-8 value will be \xe3\x81\x82 and lastly the SHIFT-JIS value will be \x82\xa0.
All of these can be directly represented as a stream of bytes with direct mappings, however converting between these seems to be a challenge.
With a language such as Python, the below could be used:
utf8_val = "あ"
unicode_val = utf8_val.decode("utf-8")
sjis_val = unicode_val.encode("shift-jis", "ignore")
I understand that ICU or Boost.Locale can achieve similar results using to_utf and from_utf functions, however why does there not seem to be any basic way to achieve this using the standard library?
It seems conversion mappings exist in Microsoft's STL (https://github.com/microsoft/STL/tree/master/stl/inc/cvt) however with no real documentation about them, it isn't clear how to go about utilizing them.
Lastly, it seems that this is achieveable through the <locale> library, however documentation does not seem to have clear answers on how. My best guess is to do with codecvt::in and codecvt::out however even here, it isn't clear on a proper approach.
Am I simply overlooking something? Or is conversion really this difficult?
edit: I have tried to do a basic char->unicode conversion using documentation from MS (here):
std::string conv(std::string str, std::string enc)
{
std::locale loc(enc.c_str());
std::string dest;
dest.resize(str.length() * 2);
char* pnewNext;
const char* porigNext;
std::mbstate_t mbstate;
auto res = std::use_facet<std::codecvt<char, char, std::mbstate_t>>
(loc).out(mbstate,
&str[0], &str[str.size()-1], porigNext,
&dest[0], &dest[dest.length()-1], pnewNext);
std::cout << "errcode : " << res << std::endl;
for(const auto& c : str)
{
std::cout << std::hex << std::setw(2) << std::setfill('0') << (uint)(unsigned char)c << std::endl;
}
for(const auto& c : dest)
{
std::cout << std::hex << std::setw(2) << std::setfill('0') << (uint)(unsigned char)c << std::endl;
}
return dest;
}
Unfortunately this still does not seem to be a correct approach:
errcode : 3 // this is std::codecvt_base::noconv
33 //these are the input bytes
44
00 //these are the output bytes
00
00
00
For anyone who may come across this, I wrote a (poor) UTF-8 to Unicode conversion algorithm. It could definitely be improved (both in performance and code size). Regardless, it successfully converts (without any real checks - be warned).
std::size_t ucsz(char c)
{
std::size_t count = 0;
for (std::size_t i = 7; i > 0; --i)
{
if (((c >> i) & 1) == 0)
{
return std::max(static_cast<int>(count), 1);
}
++count;
}
return count;
}
std::string utf8_to_unicode(std::string utf8)
{
std::string unicode;
for (std::size_t i = 0; i < utf8.length();)
{
auto length = ucsz(utf8[i]);
switch (length)
{
case 4:
{
std::uint32_t val = 0;
for (std::size_t j = 0; j < 21; ++j)
{
auto idx = std::max(0, (int)(i + (3 - (j / 6))));
val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
}
unicode.push_back(val >> 16);
unicode.push_back(val >> 8);
unicode.push_back(val >> 0);
++i;
++i;
++i;
break;
}
case 3:
{
std::uint32_t val = 0;
for (std::size_t j = 0; j < 16; ++j)
{
auto idx = std::max(0, (int)(i + (2 - (j / 6))));
val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
}
unicode.push_back(val >> 8);
unicode.push_back(val >> 0);
++i;
++i;
break;
}
case 2:
{
std::uint16_t val = 0;
for (std::size_t j = 0; j < 11; ++j)
{
auto idx = std::max(0, (int)(i + (1 - (j / 6))));
val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
}
unicode.push_back(val >> 8);
unicode.push_back(val >> 0);
++i;
break;
}
case 1:
{
unicode.push_back(utf8[i]);
break;
}
default:
{
break;
}
}
++i;
}
return unicode;
}