Fast single precision reciprocal square root in C++ on very low precision

Viewed 366

I have a line in C++

c[i] = sqrtf(a[i]);

and assembly code looks

002D11D0  vsqrtps     ymm0,ymmword ptr a (202D3380h)[eax]  

With a line

c[i] = 1.0f / sqrtf(a[i]);

i have an assembly

00E71210  vrsqrtps    ymm1,ymm0  
00E71214  vmulps      ymm0,ymm1,ymm0  
00E71218  vmulps      ymm0,ymm0,ymm1  
00E7121C  vsubps      ymm0,ymm0,ymm6  
00E71220  vmulps      ymm0,ymm0,ymm1  
00E71224  vmulps      ymm0,ymm0,ymm7

It is obviously reasonable becouse vrsqrtps is much more faster than vsqrtps. So in case of reciprocal value of square root its simply faster to call non-accurate function vrsqrtps and then do some two iteration to get more precise value.

And my question is: Is it possible to tell the compiler that additional iterations are not nessesary? So the assembly will be without additional multiplications. Error of ~1.5 * 2^-12 is fully sufficient for me, since i want to add thousands of these results where many bits of accuracy will be droped as well. I prefer a way to not inlining some assembly code into C++ code.

(after edit) Compiler comand line:

/GS /Qpar /GL /analyze- /W3 /Gy /Zc:wchar_t /Zi /Gm- /Ox /Ob2 /sdl /Fd"Release\vc141.pdb" /Zc:inline /fp:fast /D "_MBCS" /errorReport:prompt /WX- /Zc:forScope /arch:AVX2 /Gd /Oy- /Oi /MD /Fa"Release\" /EHsc /nologo /Fo"Release\" /Ot /Fp"Release\performancetest.pch" /diagnostics:classic 
1 Answers

I am afraid, there are no compiler flags to force reduced precision calculation for the reciprocal square root.

But you can easily write your own function for that using intrinsics, for example:

#include <immintrin.h>

float fast_rsqrt( float x ) {
    return _mm_cvtss_f32( _mm_rsqrt_ss( _mm_set_ss( x ) ) );
}

This will work on x86-64 platforms in Clang, GCC and MSVC compilers. The latest MSVC will produce such assembly code for it:

float fast_rsqrt(float) PROC                                ; fast_rsqrt
        rsqrtss xmm1, xmm0
        movaps  xmm0, xmm1
        ret     0
float fast_rsqrt(float) ENDP                                ; fast_rsqrt

Demo: https://gcc.godbolt.org/z/dE47M6a8x

Related