I have the following C/C++ code:
#define SIZE 2
typedef struct vec {
float data[SIZE];
} vec;
vec add(vec a, vec b) {
vec result;
for (size_t i = 0; i < SIZE; ++i) {
result.data[i] = a.data[i] + b.data[i];
}
return result;
}
I was wondering how clang would optimize this vector addition and the compiler output surprised me, as it looks quite unoptimal. This is at -O3 and with -march=skylake. (Godbolt with clang 10.1)
add(vec, vec):
vaddss xmm2, xmm0, xmm1 # res[0] = a[0] + b[0]
vmovss dword ptr [rsp - 8], xmm2 # mem[1] = res[0]
vmovshdup xmm0, xmm0 # a[0] = a[1]
vmovshdup xmm1, xmm1 # b[0] = b[1]
vaddss xmm0, xmm0, xmm1 # a[0] = a[0] + b[0]
vmovss dword ptr [rsp - 4], xmm0 # mem[0] = a[0]
vmovsd xmm0, qword ptr [rsp - 8] # xmm0 = mem[0],mem[1],zero,zero
ret
From the looks of it, a and b are stored in xmm0 and xmm1 respectively. However, only the lowest single-precision float in these registers is being used for addition. This leads to two separate additions. Why isn't vaddps used instead, which would allow for adding both values simultaneously?
The only thing I could come up with is that clang tries to preserve the higher two floats in the xmm registers. This is why I also tried increasing SIZE to 4, but now I get:
add(vec, vec):
vaddps xmm0, xmm0, xmm2
vaddps xmm1, xmm1, xmm3
vmovlhps xmm0, xmm0, xmm1
vmovaps xmmword ptr [rsp - 24], xmm0
vmovsd xmm0, qword ptr [rsp - 24]
vmovsd xmm1, qword ptr [rsp - 16]
ret
So for whatever reason, clang now doesn't even use the highest two floats and spreads the vectors between xmm0 to xmm3. An xmm register is 128 bits large, so it should be able to fit all four floats. Then this code would be much simpler and only a single addition would be necessary.
(See Compiler Explorer)