SpiceQA
Questions
Tags
Users
Badges
sse
223 Questions
Newest
Active
Unanswered
Frequent
More
Score
View
Card
Compact
_mm256_fmadd_ps is slower than _mm256_mul_ps + _mm256_add_ps?
user_1422624
0
•
asked Feb 18, 2021
2
1
1220
avx
micro-optimization
simd
sse
gcc
How to best emulate the logical meaning of _mm_slli_si128 (128-bit bit-shift), not _mm_bslli_si128
user_4308438
0
•
asked Feb 7, 2021
4
2
706
intrinsics
sse2
simd
sse
c
Why VS C++ 2017 compiler use SSE optimization only if iterated pointers are not stored in structure?
user_8132066
0
•
asked Jan 27, 2021
2
0
123
visual-c++-2017
compiler-optimization
sse
visual-c++
c++
Which are the use case of punpcklbw (interleave in MMX/SSE/AVX)?
user_1447389
0
•
asked Jan 25, 2021
2
1
278
memset
disassembly
sse
assembly
compression
x86 SIMD instructions 16 byte alignment in assembly (Without C intrinsics)
user_14152996
0
•
asked Jan 10, 2021
3
1
685
x86-64
simd
memory-alignment
sse
assembly
Expand the lower two 32-bit floats of an xmm register to the whole xmm register
user_14152996
0
•
asked Jan 9, 2021
2
1
176
sse
x86
assembly
On x86-64, is the “movnti” or "movntdq" instruction atomic when system crash?
user_14883676
0
•
asked Jan 4, 2021
5
1
521
persistent-memory
x86-64
atomic
sse
cpu-architecture
What is the most efficient way to do unsigned 64 bit comparison on SSE2?
user_3900123
0
•
asked Dec 24, 2020
2
2
438
sse2
simd
sse
assembly
How could the movdqa/movdqu instruction write an incorrect value to memory?
user_3807417
0
•
asked Dec 18, 2020
2
1
364
xcode-debugger
sse
xcode
x86
assembly
Multiply-add vectorization slower with AVX than with SSE
user_1094044
0
•
asked Dec 14, 2020
5
2
303
avx
sse
optimization
c++
performance
Prev
Prev
7
8
9
(current)
10
11
Next
Next
Hot Questions