Is clang really adding vectors optimally in this C/C++ example?

Viewed 116

I have the following C/C++ code:

#define SIZE 2

typedef struct vec {
    float data[SIZE];
} vec;

vec add(vec a, vec b) {
    vec result;
    for (size_t i = 0; i < SIZE; ++i) {
        result.data[i] = a.data[i] + b.data[i];
    }
    return result;
}

I was wondering how clang would optimize this vector addition and the compiler output surprised me, as it looks quite unoptimal. This is at -O3 and with -march=skylake. (Godbolt with clang 10.1)

add(vec, vec):                          
        vaddss  xmm2, xmm0, xmm1          # res[0] = a[0] + b[0]
        vmovss  dword ptr [rsp - 8], xmm2 # mem[1] = res[0]
        vmovshdup       xmm0, xmm0        # a[0] = a[1]
        vmovshdup       xmm1, xmm1        # b[0] = b[1]
        vaddss  xmm0, xmm0, xmm1          # a[0] = a[0] + b[0]
        vmovss  dword ptr [rsp - 4], xmm0 # mem[0] = a[0]
        vmovsd  xmm0, qword ptr [rsp - 8] # xmm0 = mem[0],mem[1],zero,zero
        ret

From the looks of it, a and b are stored in xmm0 and xmm1 respectively. However, only the lowest single-precision float in these registers is being used for addition. This leads to two separate additions. Why isn't vaddps used instead, which would allow for adding both values simultaneously?

The only thing I could come up with is that clang tries to preserve the higher two floats in the xmm registers. This is why I also tried increasing SIZE to 4, but now I get:

add(vec, vec):      
        vaddps  xmm0, xmm0, xmm2
        vaddps  xmm1, xmm1, xmm3
        vmovlhps        xmm0, xmm0, xmm1 
        vmovaps xmmword ptr [rsp - 24], xmm0
        vmovsd  xmm0, qword ptr [rsp - 24] 
        vmovsd  xmm1, qword ptr [rsp - 16]
        ret

So for whatever reason, clang now doesn't even use the highest two floats and spreads the vectors between xmm0 to xmm3. An xmm register is 128 bits large, so it should be able to fit all four floats. Then this code would be much simpler and only a single addition would be necessary.

(See Compiler Explorer)

0 Answers
Related