Why is the first packed data in struct little endian, but the rest is big endian?

Viewed 369
import struct
port = 1331
fragments = [1,2,3,4]
flags = bytes([64])
name = "Hello World"

data = struct.pack('HcHH', port, flags, len(fragments), len(name))

print(int.from_bytes(data[3:5], byteorder='big'))
print(int.from_bytes(data[5:7], byteorder='big'))
print(int.from_bytes(data[0:2], byteorder='little'))

When I print them like this, they come out correctly. It seems port is in little endian, while len(fragments) and len(name) are in big endian. If I also do big endian on the port, it gets the wrong value.

So why does struct behave like this? Or am I missing something?

2 Answers

There is some funny alignment taking place because of the 'c' in the middle of 'H'. You can see it with calcsize:

>>> struct.calcsize('HcHH')
8
>>> struct.calcsize('HHHc')
7

So your data is not aligned as you thought. The correct unpacking is:

print(int.from_bytes(data[4:6], byteorder='little'))
# 4
print(int.from_bytes(data[6:], byteorder='little'))
# 11

It turns out that by chance, the added byte of the 'c' is '\x00', and made your byte-chain correct in big-endian:

>>> data
b'3\x05@\x00\x04\x00\x0b\x00'
        ^^^^
        this is the intruder

By default, your call to pack is equivalent to the following:

struct.pack('@HcHH', port, flags, len(fragments), len(name))

The result looks like this (printed with '.'.join(f'{x:02X} for x in data')):

33.05.40.00.04.00.0B.00
 0  1  2  3  4  5  6  7

The number 4 is encoded in bytes 4 and 5, in little endian, and 11 is encoded in bytes 6 and 7. Byte 3 is a padding byte, inserted by pack to properly align the following shorts on an even boundary.

Per the docs:

Note By default, the result of packing a given C struct includes pad bytes in order to maintain proper alignment for the C types involved; similarly, alignment is taken into account when unpacking. This behavior is chosen so that the bytes of a packed struct correspond exactly to the layout in memory of the corresponding C struct. To handle platform-independent data formats or omit implicit pad bytes, use standard size and alignment instead of native size and alignment: see Byte Order, Size, and Alignment for details.

To remove the alignment byte and justify your assumptions about the positions of the bytes while keeping native byte order, use

struct.pack('=HcHH', port, flags, len(fragments), len(name))

You can also use a fixed byte order by using < or > as the prefix.

The "correct" solution is to use unpack to get your numbers back, so you don't have to worry about endianness, padding or anything else, really.

Related