Are there better AVX instructions to move data from 3 ymm registers?

Viewed 337

I have three ymm registers -- ymm4, ymm5 and ymm6 -- packed with double precision (qword) floats:

ymm4:   73  144 168 41
ymm5:   144 348 26  144
ymm6:   732 83  144 852

I want to write each column of the matrix above. For example:

-- extract ymm4[63:0] and insert it at ymm0[63:0]
-- extract ymm5[63:0] and insert it at ymm0[127:64]
-- extract ymm6[63:0] and insert it at ymm0[191:128]

so that ymm0 reads 73, 144, 732.

So far I have used:

mov rax,4
kmovq k6,rax
vpxor ymm1,ymm1
VEXPANDPD ymm1{k6}{z},ymm6

That causes ymm1 to read [ 0 0 732 ], so I've accomplished the first step because 732 is the element at [63:0] in ymm6.

For ymm4 and ymm5 I use vblendpd:

vblendpd ymm0,ymm1,ymm4,1

That causes ymm0 to read [ 73 0 732 ], so I've accomplished the second step because 73 is the element at [63:0] in ymm4.

Now I need to put ymm5[63:0] at ymm0[127:64]:

vblendpd ymm0,ymm0,ymm5,2

That causes ymm0 to read [ 73 144 732 ], so now I am finished with the first column [63:0].

But now I need to do the same thing with columns 2, 3 and 4 in the ymm registers. Before I add more instructions, is this the most efficient way to do what I described? Is there another, more effective way?

I have investigated unpckhpd (https://www.felixcloutier.com/x86/unpckhpd), vblendpd (https://www.felixcloutier.com/x86/blendpd, and vshufpd (https://www.felixcloutier.com/x86/shufpd), and what I show above seems like the best solution but it's a lot of instructions, and the encodings shown in the docs for the imm8 value are somewhat opaque. Is there a better way to extract the corresponding columns of three ymm registers?

1 Answers

Let's name the matrix elements like this:

YMM0 = [A,B,C,D]
YMM1 = [E,F,G,H]
YMM2 = [I,J,K,L]

Eventually, you want a result like this, where * indicates a “don't care.”

YMM0 = [A,E,I,*]
YMM1 = [B,F,J,*]
YMM2 = [C,G,K,*]
YMM3 = [D,H,K,*]

To achieve this, we extend the matrix to 4×4 (imagine another row of just [*,*,*,*]) and then transpose the matrix. This is done in two steps: first, each 2×2 submatrix is transposed. Then, the top left and bottom right matrices are exchanged:

[A,B,C,D]       [A,E,C,G]       [A,E,I,*]
[E,F,G,H]  --\  [B,F,D,H]  --\  [B,F,J,*]
[I,J,K,L]  --/  [I,*,K,*]  --/  [C,G,K,*]
[*,*,*,*]       [J,*,L,*]       [D,H,L,*]

For the first step in ymm0 and ymm1, we use a pair of unpack instructions:

vunpcklpd %ymm1, %ymm0, %ymm4         // YMM4 = [A,E,C,G]
vunpckhpd %ymm1, %ymm0, %ymm5         // YMM5 = [B,F,D,H]

Row 3 stays in ymm2 for the moment as it doesn't need to be changed. Row 4 is obtained by unpacking ymm2 with itself:

vunpckhpd %ymm2, %ymm2, %ymm6         // YMM5 = [J,*,L,*]

The second step is achieved by blending and swapping lanes twice:

vblendpd $0xa, %ymm2, %ymm4, %ymm0    // YMM0 = [A,E,I,*]
vblendpd $0xa, %ymm6, %ymm5, %ymm1    // YMM1 = [B,F,J,*]
vperm2f128 $0x31, %ymm2, %ymm4, %ymm2 // YMM2 = [C,G,K,*]
vperm2f128 $0x31, %ymm6, %ymm5, %ymm3 // YMM3 = [D,H,L,*]

This achieves the desired permutation in 7 instructions.

Note that as none of these instructions require AVX2, this code will run on a Sandy Bridge processor with just AVX.

Related