Why Clang allocate std::complex so fast?

Viewed 151

I have tested the memory allocation performance for std::complex<double> array using GCC, Clang, and Intel C++ compilers. I found that there is a dramatic performance difference between these compilers. Clang is several orders of magnitude faster than other compilers.

Can anyone figures out the reason? Thanks so much! The test code, test environment, and test result details are attached below.

Attachments

Here is the performance test code:

#include <iostream>
#include <complex>

#include <sys/time.h>
#include <stdlib.h>


double GetWallTime(void) {
  struct timeval time;
  if (gettimeofday(&time, NULL)) { return 0; }
  return (double)time.tv_sec + (double)time.tv_usec * .000001;
}


int main(int argc, char *argv[]) {
  auto size = atol(argv[1]);

  auto time_now = GetWallTime();
  auto double_array = new double [size];
  std::cout << "allocate double[" << size << "]: use "
            << GetWallTime() - time_now   << " sec" << std::endl;
  delete[] double_array;

  time_now = GetWallTime();
  auto c99cplx_array = new double _Complex [size];
  std::cout << "allocate double _Complex[" << size << "]: use "
            << GetWallTime() - time_now  << " sec" << std::endl;
  delete[] c99cplx_array;

  time_now = GetWallTime();
  auto stdcplx_array = new std::complex<double> [size];
  std::cout << "allocate std::complex<double>[" << size << "]: use "
            << GetWallTime() - time_now         << " sec" << std::endl;
  delete[] stdcplx_array;
  return 0;
}

On a Linux machine with Xeon E5-2680 v2 *2 CPUs and 64G DDR3 memory, I compiler this code using these three compilers and run the test. I obtain the following results:

Use g++ 7.5.0 with --std=c++11 -O3 flags:

allocate double[10000000]: use 5.6982e-05 sec
allocate double _Complex[10000000]: use 1.71661e-05 sec
allocate std::complex<double>[10000000]: use 0.0911679 sec

Use icpc 19.1.0.166 with --std=c++11 -O3 flags:

allocate double[10000000]: use 7.41482e-05 sec
allocate double _Complex[10000000]: use 1.69277e-05 sec
allocate std::complex<double>[10000000]: use 0.087034 sec

Use clang++ with the version

> clang++ --version
Intel(R) oneAPI DPC++ Compiler 2021.1-beta03 (2019.10.0.1121)
Target: x86_64-unknown-linux-gnu
Thread model: posix

with --std=c++11 -O3 flags:

allocate double[10000000]: use 4.19617e-05 sec
allocate double _Complex[10000000]: use 9.53674e-07 sec
allocate std::complex<double>[10000000]: use 1.19209e-06 sec
1 Answers

Because your test is a terrible one. Clang has simply removed the new/delete completely because it's pointless. You are effectively profiling this:

#include <complex>
int main() {

    auto ptr = new std::complex<double>[1000];
    delete [] ptr;
    return 0;
}

Which clang has correctly reduced to:

main:                                   # @main
        xor     eax, eax
        ret

https://godbolt.org/z/4JjXfh

gcc on the other hand allocates and frees the memory. You aren't testing allocation speed, you are testing a nothing burger. Don't micro benchmark like this, instead use a profiler on actual code.

If you want to profile this, do:

volatile auto stdcplx_array = new std::complex<double> [size];

Which will provide an accurate benchmark for each compiler (not that your benchmark is worth doing - it will just reduce to a new + memset + delete).

Related