Why does the UTF-8 encoding have the concept of overlongs?

Viewed 86

This is not a question along the lines of, "what is an overlong?" or, "what should I do with overlongs?" as I understand what an overlong is and I understand how they should be handled. This is a question possibly about history and possibly about some limitation I don't understand.

In the UTF-8 encoding scheme you could encode the same binary sequence in multiple ways, for example:

00101010, 11000000 10101010, 11100000 10000000 10101010, and 11110000 10000000 10000000 10101010

All technically decode to the same binary sequence 101010, which all represent the number 42 just with a variable amount of leading zeros. Of course the only valid encoding in UTF-8 is the shortest. The rest are called overlongs and are strictly not valid UTF-8.

But

It seems like this is both:

  1. Space wasteful
  2. Parser complicating

If instead each multi-byte sequence was given a starting integer offset then it seems there would be:

  1. No such thing as an overlong
  2. Simpler logic for parsers to implement
  3. More available numbers to represent characters

The offsets would simply be the next possible integer to represent.

byte length offset usable bits
1 0 7
2 2^7 = 128 11
3 2^11 = 2048 16
4 2^16 = 65536 21

Then the sequences listed above would all have different values:

  • 00101010 = 42
  • 11000000 10101010 = 128 + 42
  • 11100000 10000000 10101010 = 2048 + 42
  • 11110000 10000000 10000000 10101010 = 65536 + 42

and the maximum UTF-8 value would go from 2^21 to 2^21 + 65536.

Is there a technical or historical reason that this isn't the case?

1 Answers

I think it is just for simplicity (in origin). Your proposal is sensible, and UTF-16 use it (so adding a constant to the bits given by surrogates).

But does it help? As you may see, you may get very little efficiency: check the character which can be shortened with your proposal: not really the most used characters, so not much about compressing text. And UTF-8 with self-synchronizing is also not designed to be the shortest sequence.

As you see in comments, original UTF-8 allowed all UCS characters, so 31-bits. Only later (and because of UTF-16 limitation), UCS and Unicode decided that the maximum characters should be U+10FFFF, so limiting UTF-8 to 4 bytes.

Note: now the implementation is not so simple, because one should check that there are not overlong sequences (it is a security risk), not using values above maximum allowed codepoint, and also no values in the surrogate sequence. So now simplicity is not really true, and with your proposal, we would skip the first check (and possibly the most nasty one).

Note: the overlong sequences help to encode \0, which it is a valid character on a Unicode string, but used as string termination on C (and so on many languages and API). I suspect this may also be a reason, OTOH overlong sequences makes an invalid UTF-8 string (it is MUTF-8). But I never saw elements which could confirm this.

Related