x86 SIMD instructions 16 byte alignment in assembly (Without C intrinsics)

Viewed 685

Let's say that I have an array of 8 byte elements of unknown length from memory passed to my assembly function. I want to do some 128 bit SIMD operations (up to SSE4) on it. It is better that the memory is 16 byte aligned. So I would check if the array is aligned and then depending on that use movaps or movups.

I know you can check 16-byte alignment with:

test dil, 0xf        ; rdi stores address of array

If it isn't 16-byte aligned, is it good or useful to also check if it's 8-byte aligned, which would mean that it's an odd multiple of 8?

test dil, 0x7           ; ZF=1 here after rdi&0xf !=0 implies rdi%16 == 8

And if that is true then should I do an extra step on the first element of array and then movaps to load the array elements? And otherwise should I just use unaligned operations like movups?

Does it work like this?

1 Answers

If your arrays are usually aligned by 16, it's probably best not to do even more checking to look for the odd-start case, just use your unaligned version unless it's a lot worse for some reason.

However, if they're usually aligned by 8 (but unknown whether they're aligned by 16), then you may be able to get away with only checking for alignment by 8 and branchlessly handling the maybe-unaligned first iteration for the aligned case, see below. (Otherwise just fall back to your fully unaligned case.)

If overlap isn't a problem (e.g. c[] = a[]+b[], or a memset-like store or whatever), a good technique is to always do a first vector with unaligned load/store, then advance to the first aligned vector (add rdi, 16 / and rdi, -16). If the input was aligned, this won't overlap. Otherwise, it partially overlaps and the store buffer + L1d cache handle it efficiently.

This keeps the cost minimal for the aligned case, and avoids the chance of branch mispredicts.

Rounding a pointer up/down to an alignment boundary is cheap, just an and, but you do have the code-size cost of peeling a whole copy of the loop body. So it's not totally free as far as startup overhead, but at least this kind of startup overhead can overlap with a cache miss from the data.


But note that a lot of SIMD functions have multiple pointer inputs that can be misaligned relative to each other. In that case, the standard advice is to align the output and keep using movups for inputs. Although if the front-end is the bottleneck, you might choose to reach an alignment boundary for the input so you can fold a memory source operand into an ALU instruction like xorps xmm0, [rdi] and use movups store.

But if anything other than the front-end, e.g. cache or memory throughput, are a bottleneck, then you more often want to align the destination. Intel's optimization manual has some advice about this. Some of the reason is that load throughput is typically 2x store throughput (until IceLake), so the load hardware can more readily absorb the extra work for split loads. Also, storing a full cache line with fewer stores can help reduce cases where a line gets evicted (written back) but then you store to it again and it has to get fetched + dirtied and eventually written-back again, instead of just fetched.

Related