Why must character encoding conversion be so difficult in c++?

Viewed 309

As a foreword, I don't doubt that my knowledge of Unicode and character encoding is incomplete. Sorry for any misrepresentations.

I have recently been developing an application that will deal with incoming data in various encodings. These are, for example, UTF-8, ASCII, SHIFT-JIS, etc.

Now my understanding is that for a given encoding, a 'mapping exists' between it and a given Unicode value. Take for example the character 'あ'.

The Unicode value will be U+3042, the UTF-8 value will be \xe3\x81\x82 and lastly the SHIFT-JIS value will be \x82\xa0.

All of these can be directly represented as a stream of bytes with direct mappings, however converting between these seems to be a challenge.

With a language such as Python, the below could be used:

utf8_val = "あ"
unicode_val = utf8_val.decode("utf-8")
sjis_val = unicode_val.encode("shift-jis", "ignore")

I understand that ICU or Boost.Locale can achieve similar results using to_utf and from_utf functions, however why does there not seem to be any basic way to achieve this using the standard library?

It seems conversion mappings exist in Microsoft's STL (https://github.com/microsoft/STL/tree/master/stl/inc/cvt) however with no real documentation about them, it isn't clear how to go about utilizing them.

Lastly, it seems that this is achieveable through the <locale> library, however documentation does not seem to have clear answers on how. My best guess is to do with codecvt::in and codecvt::out however even here, it isn't clear on a proper approach.

Am I simply overlooking something? Or is conversion really this difficult?

edit: I have tried to do a basic char->unicode conversion using documentation from MS (here):

std::string conv(std::string str, std::string enc)
{
    std::locale loc(enc.c_str());
    std::string dest;
    dest.resize(str.length() * 2);

    char* pnewNext;
    const char* porigNext;
    std::mbstate_t mbstate;

    auto res = std::use_facet<std::codecvt<char, char, std::mbstate_t>>
               (loc).out(mbstate,
                         &str[0], &str[str.size()-1], porigNext,
                         &dest[0], &dest[dest.length()-1], pnewNext);

    std::cout << "errcode : " << res << std::endl;
    for(const auto& c : str)
    {
        std::cout << std::hex << std::setw(2) << std::setfill('0') << (uint)(unsigned char)c << std::endl;
    }
    for(const auto& c : dest)
    {
        std::cout << std::hex << std::setw(2) << std::setfill('0') << (uint)(unsigned char)c << std::endl;
    }
    return dest;
}

Unfortunately this still does not seem to be a correct approach:

errcode : 3 // this is std::codecvt_base::noconv

33 //these are the input bytes
44

00 //these are the output bytes
00
00
00

For anyone who may come across this, I wrote a (poor) UTF-8 to Unicode conversion algorithm. It could definitely be improved (both in performance and code size). Regardless, it successfully converts (without any real checks - be warned).

std::size_t ucsz(char c)
{
    std::size_t count = 0;
    for (std::size_t i = 7; i > 0; --i)
    {
        if (((c >> i) & 1) == 0)
        {
            return std::max(static_cast<int>(count), 1);
        }
        ++count;
    }
    return count;
}

std::string utf8_to_unicode(std::string utf8)
{
    std::string unicode;
    for (std::size_t i = 0; i < utf8.length();)
    {
        auto length = ucsz(utf8[i]);
        switch (length)
        {
        case 4:
        {
            std::uint32_t val = 0;
            for (std::size_t j = 0; j < 21; ++j)
            {
                auto idx = std::max(0, (int)(i + (3 - (j / 6))));
                val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
            }
            unicode.push_back(val >> 16);
            unicode.push_back(val >> 8);
            unicode.push_back(val >> 0);
            ++i;
            ++i;
            ++i;
            break;
        }
        case 3:
        {
            std::uint32_t val = 0;
            for (std::size_t j = 0; j < 16; ++j)
            {
                auto idx = std::max(0, (int)(i + (2 - (j / 6))));
                val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
            }
            unicode.push_back(val >> 8);
            unicode.push_back(val >> 0);
            ++i;
            ++i;
            break;
        }
        case 2:
        {
            std::uint16_t val = 0;
            for (std::size_t j = 0; j < 11; ++j)
            {
                auto idx = std::max(0, (int)(i + (1 - (j / 6))));
                val |= ((utf8[idx] >> (j % 6)) & 1U) ? (1UL << j) : 0;
            }
            unicode.push_back(val >> 8);
            unicode.push_back(val >> 0);
            ++i;
            break;
        }
        case 1:
        {
            unicode.push_back(utf8[i]);
            break;
        }
        default:
        {
            break;
        }
        }
        ++i;
    }
    return unicode;
}
0 Answers
Related