Why does BigInteger parse "80"(hex) to two bytes?

Viewed 151

I want to convert a hex-string to a byte-array. I thought using BigInteger is a good idea. But for values greater than 7F it produces unexpected results.

My code:

    var bytes = new BigInteger("80", 16).toByteArray();
    for (var b : bytes) System.out.println(b);

It outputs:

0
-128

Why does this produce two bytes?

I would have expected 00 to FF to produce one byte, 0100 to FFFF to produce two bytes, and so forth.


Side note: The first byte seems to actually matter:

new BigInteger(new byte[]{      (byte)0x80}); // produces -128 (negative!)
new BigInteger(new byte[]{   0, (byte)0x80}); // produces  128
new BigInteger(new byte[]{0, 0, (byte)0x80}); // produces  128
4 Answers

Because BigInteger is signed.

You specified "80" in hex, and you didn't specify that it is negative; therefore the highest bit (in two's complement) must be zero. If you try to represent 80 in one byte, then the top bit is 1, so it would be negative.

If you try new BigInteger("-80", 16).toByteArray() then you get one byte, with the value -128.

the toByteArray() method returns a byte array containing the two's-complement representation of this BigInteger. The byte array will be in big-endian byte-order

the method internally looks like:

public byte[] toByteArray() {
    int byteLen = bitLength()/8 + 1;
    byte[] byteArray = new byte[byteLen];
    for (int i=byteLen-1, bytesCopied=4, nextInt=0, intIndex=0; i >= 0; i--) {
        if (bytesCopied == 4) {
            nextInt = getInt(intIndex++);
            bytesCopied = 1;
        } else {
            nextInt >>>= 8;
            bytesCopied++;
        }
        byteArray[i] = (byte)nextInt;
    }
    return byteArray;
}

so as you can see

int byteLen = bitLength()/8 + 1;

The documentation of toByteArray() says

Returns a byte array containing the two's-complement representation of this BigInteger.

Two’s complement stores the sign bit in the highest order bit. As a consequence, signed, positive values from 0x0–0x7F (that is, 0000 0000–0111 1111 in binary) only require a single byte to store, but values greater than that require a second byte, since they’d otherwise denote a negative value. In particular, 1000 0000 (which, when unsigned, can be written as 0x80) corresponds to the value −0x80, not +0x80, in two’s complement.

Thanks for all Your contributions, comments and answers, finally I understand what is going on.

Numbers from 0 to Hex 0x7F (decimal 127) work like this:

First Bit=sign (0=+,1=-)
|    Other bits=number
v    vvv vvvv
0    111 1111    = 0x7F = 127

Numbers from 128 to Hex 0x7FFF (decimal 32.767) work like this:

First Bit (of first byte!)=sign
|    Other bits and other bytes=number
v    vvv vvvv vvvv vvvv
0    000 0000 1000 0000    = 0x80 = 128

Long story short:

  • Only the first bit of the first byte is determining the sign
  • All other bits of all other bytes determine the absolute value
    • That is 7 bits of the first byte
    • And 8 bits of every other byte
Related