Can an empty struct (or struct padding) have an arbitrary value?

Viewed 114

Let's say I have this code that has UB:

union Flag {
    constexpr Flag() : empty{} {}

    struct {} empty;
    bool value;
};

static Flag flag;

int main() {
    return flag.value;
}

where the UB is accessing value when it’s not the active member of the union Flag.

Currently, UBSan will not catch this error because (as I understand it) UBSan does not have a way to check the last written member of a union. For this particular case, I think UBSan could catch some UB going on here indirectly via the same check for non-true/false values for type bool. If the byte of an empty struct is considered “uninitialized” or “can have any arbitrary value”, then the compiler could legally set the arbitrary byte of this empty struct to to any non-zero/one value, and then UBSan would be able to catch the load of an invalid value for bool.

What I'd like to know is: Is it semantically permissible to have the nominal byte of empty structs–and more generally, any padding bytes in all structs–be initialized to any non-zero pattern?

2 Answers

Is it semantically permissible to have the nominal byte of empty structs–and more generally, any padding bytes in all structs–be initialized to any non-zero pattern?

Any bytes that are not part of the value representation of a type are fair game, as far as the compiler is concerned. Well, to some degree.

You can validly memcpy into such bytes, but this is only valid if the source data comes (directly or indirectly) from an existing object of that type. This is from [basic.types]/2&3. So within the object model, the user doesn't get to just put whatever in that storage.

As such, for code that's living within the C++ object model, an implementation is allowed to play around with the contents of padding bytes.

C++20's implicit object creation rules make this rather more difficult, as it allows uninitialized objects to be manifested in storage that already has bytes in it. These manifestations aren't typically associated with code like placement-new, so it would be very difficult for UBSan to initialize such a thing.

The union you've shown is implicit lifetime (due to the trivial copy/move constructors), so users can play games with such things.

Certain parts of the Standard only really make sense if one accepts the possibility that certain union objects may sometimes be eligible to be read as any type until the next time they are written. Most notably, if a union object containing only trivial types is written using fread, memcpy, or other such means, from a byte source that has no identifiable association with anything in the union, there would often be no way a compiler could know which union member was being initialized, so reading any member for which the byte sequence would be a valid representation would be valid.

A compiler would not have to zero-fill an empty structure within a union, and a conforming but capricious compiler could decide to zero-fill structures only when their bitwise representation is observed using character types, but if code uses a character type to observe the initial byte pattern of a union that is never written and also reads that storage using another type like bool, a compiler would be required to either treat the bit pattern as valid in the latter read, or else have ensured that the character-type access reported a bit pattern that's not valid for that type. The easiest way for a compiler to uphold that requirement is to simply treat the access as valid.

Related