I have a 32-bit C++ DSP audio processing project on Analog Devices Sharc DSP processor and need to move it to 64-bit processing that has been now for some time available for embedded use-cases with ARM AArch64.
I am considering two alternatives:
- either to use my own custom implementation of FIR and IIR filtering or
- to go for some library functions optimized for AArch64 and Neon.
I have quite CPU-intensive processing on top of 64-bit precision. I also need to gain much more processing power as currently Sharc performance is also a bottleneck. The IIR and FIR functions should provide 64-bit real-time, block-based signal processing.
My target platform is Raspberry Pi, 3B+ maybe 4. The sort of functions I need is provided e.g. in the CMSIS library as arm_biquad_cascade_df2T_f64() (it actually works together with a supplementary init function that implements the state array needed to process data in a block-based manner). And the library funcs seems to work with 64-bit. But I have doubts if they are suitable and optimized for AArch64 as generally CMSIS is labeled 32-bit, Ne10 similarly.
I am exploring the custom code path, my questions are:
- what kind of Neon and AArch64 specific optimizations are possible
- which magnitude of performance improvement can be expected compared to plain C implementation of block-based biquad function
Or maybe it is enough when it is left to compiler optimization and use of Neon?