SpiceQA
Questions
Tags
Users
Badges
neon
50 Questions
Newest
Active
Unanswered
Frequent
More
Score
View
Card
Compact
What is the most efficient way to handle integer multiplication overflow with saturation with ARM Neon intrinsics?
user_10576494
0
•
asked Jan 6, 2022
2
1
211
neon
intrinsics
saturation-arithmetic
arm
simd
aarch64 xtn2 clearing lower half
user_8582383
0
•
asked Dec 12, 2021
3
1
95
armv8
arm64
neon
simd
assembly
Why doesn’t Clang use vcnt for __builtin_popcountll on AArch32?
user_1813349
0
•
asked Nov 17, 2021
3
1
345
neon
arm
micro-optimization
bit-manipulation
clang
Loop takes more cycles to execute than expected in an ARM Cortex-A72 CPU
user_523079
0
•
asked Nov 5, 2021
6
3
356
neon
arm
assembly
optimization
performance
NEON assembly code requires more cycles on Cortex-A72 vs Cortex-A53
user_3983330
0
•
asked Oct 26, 2021
2
1
324
raspberry-pi
arm64
neon
arm
assembly
Neon equivalent of mm_madd_epi16 and mm_maddubs_epi16
user_266836
0
•
asked Oct 21, 2021
2
1
158
neon
arm
sse
c
Why does gcc, with -O3, unnecessarily clear a local ARM NEON array?
user_523079
0
•
asked Oct 7, 2021
6
2
196
compiler-bug
arm64
neon
gcc
c
Is there a way to auto-vectorize while using reciprocals and reciprocal square roots?
user_16745102
0
•
asked Aug 24, 2021
2
0
133
auto-vectorization
avx
neon
sse
c++
What doest `vaddhn_high_s16` actually do?
user_2999096
0
•
asked Jul 18, 2021
3
1
92
arm64
neon
intrinsics
simd
c++
Why does FADDP D-form have higher throughput than FADDP Q-form on the Cortex-A72
user_3018243
0
•
asked Mar 29, 2021
2
0
188
micro-architecture
arm64
neon
simd
cpu-architecture
Prev
Prev
1
2
(current)
3
4
5
Next
Next
Hot Questions