TL;DR: Would you expect atomic_ref<3 bytes struct> to be non-lock-free, or to achieve lock-free by using CAS or LL/SC on bigger memory amount, if possible?
It is continuation of atomic_ref when external underlying type is not aligned as requested question.
Probably too big for a comment, and not enough related to edit original question.
I think that it may be possible to implement lock-free atomic_ref<T> for non natural atomic size. That is non power of two size.
Such atomic_ref<T> could access aligned memory as wider type. Most of the time it would have to fallback to compare exchange. Still it will be atomic.
It would correspond cases of T, when an implementation of atomic<T> would pad the passed type to nearest atomic size.
I think that P0091 explicitly allows to skip this, and make atomic_ref<T> only for natural atomic sizes:
Note: Whether an implementation of atomic is lock free, does not necessarily constrain whether the corresponding implementation of atomic_ref is lock free. --end note
But what is expected?
(I understand that compare-exchange is not as efficient as normal store or exchange, still I assume it can be more efficient than lock-based, and with such implementation of atomic_ref load may be even implemented as normal load)
Example:
struct S { char a, b, c; };
std::cout << std::boolalpha << std::atomic<S>::is_always_lock_free << '\n'; // I expect true
std::cout << std::boolalpha << std::atomic_ref<S>::is_always_lock_free << '\n'; // What should I expect?