Does Fortran array syntax influence speed or there is a compiler optimization fault?

Viewed 98

Let's have a double precision multidimensional array which needs to be initialized.

allocatable A(:,:,:)
! some code in which N1,N2,N3 are defined
allocate (A(N1,N2,N3))

We can initialize A in three syntax-wise ways:

A=0.d0
A(:,:,:)=0.0d0
A(1:N1,1:N2,1:N3)=0.0d0

Which one would give the fastest execution time? (to account for the variability of the execution time of the same code, neglect difference less than, say 20%)

Comments so far indicate that the execution time is not expected to be dependent on the syntax. However this test problem shows differently:

      program MatrixInit
      implicit double precision(A-H,O-Z)
      allocatable A(:,:,:)  

! writing file header      
      write(44,'(A,A)')'  N1 N2  N3       niter        tt1         tt2       ', &
                       ' tt3        tt4         tt2/tt4    tt3/tt4'
      
! Set the small dimensions N1 and N2 
! To hold, say, a stress tensor or a rotational matrix
      N1=3; N2=3    
         
! Set several examples with big N3 (say number of nodes)
      do i3=7,16
        N3=2**i3
         
        allocate(A(N1,N2,N3))
         
! Set the number of repetitions so that the times for the examples to be similar
        niter = (2**20)/(N1*N2*N3)*(2**12)
        
! NOTE: In the repetitions loop a duumy assignment is used so that the 
! initialization is not optimized to be out of the loop

! t1 convert matrix to vector - This the fastest
        call cpu_time(t0)
          do iter=1,niter
            A(1,1,1)=iter   
            call INIT(A,N1*N2*N3)
          enddo 
        call cpu_time(t1)

! t2: A=0.0d0
          do iter=1,niter
            A(1,1,1)=iter
            A=0.d0
          enddo 
        call cpu_time(t2)
      
! t3: A(:,:)=0.0d0
          do iter=1,niter
            A(1,1,1)=iter
            A(:,:,:)=0.d0
          enddo 
        call cpu_time(t3)

! t4: A(1:N2,1:N3)=0.0d0
          do iter=1,niter
            A(1,1,1)=iter
            A(1:N1,1:N2,1:N3)=0.d0
          enddo 
        call cpu_time(t4)
! Times
        tt1=t1-t0   ! Time for vector initialization
        tt2=t2-t1   ! Time for A=0.d0
        tt3=t3-t2   ! Time for A(:,:,:)=0.d0
        tt4=t4-t3   ! Time for A(1:N1,1:N2,1:N3)
      
        write(44,9)  N1,N2,N3,  niter, tt1,tt2,tt3,tt4, tt2/tt4,tt3/tt4 
       
        deallocate (A) 
      enddo
9     format(2i3,i6, 2x,i9, 6x, 4G11.4, 3x, 2G11.4)      
      end
      

      subroutine INIT(A,N)
      implicit double precision(A-H,O-Z)
      dimension A(N)
        A=0.0D0 
      return
      end

Compilation is done with Intel Fortran with Option /O2. Below is the output from the WRITE(44,...) when the program is compiled and run on two computers:

On DELL Laptop: Intel(R) Core(TM) i7-3740QM CPU @ 2.70GHz 2.70 GHz, RAM 16.0GB

Intel(R) Visual Fortran Intel(R) 64 Compiler for applications running on IA-32, Version 18.0.0.124

  N1 N2  N3       niter        tt1         tt2        tt3        tt4         tt2/tt4    tt3/tt4
  3  3   128    3727360       0.6406      6.031      7.297      1.406         4.289      5.189    
  3  3   256    1863680       0.7188      6.266      7.219      1.234         5.076      5.848    
  3  3   512     929792       0.8906      5.984      7.172      1.328         4.506      5.400    
  3  3  1024     462848       0.9219      6.000      7.250      1.359         4.414      5.333    
  3  3  2048     229376       0.9531      5.875      7.109      1.312         4.476      5.417    
  3  3  4096     114688       0.9688      5.953      7.047      1.391         4.281      5.067    
  3  3  8192      57344        1.297      6.109      7.172      1.500         4.073      4.781    
  3  3 16384      28672        1.312      5.906      7.422      1.438         4.109      5.163    
  3  3 32768      12288        1.125      5.047      6.062      1.266         3.988      4.790    
  3  3 65536       4096       0.7656      3.453      4.047     0.8906         3.877      4.544    

DELL Desktop: Intel(R) Core(TM) i7-6700 CPU @ 3.40GHz 3.41 GHz, RAM 32.0GB
Intel(R) Visual Fortran Intel(R) 64 Compiler for applications running on Intel(R) 64, Version 19.1.3.311

  N1 N2  N3       niter        tt1         tt2        tt3        tt4         tt2/tt4    tt3/tt4
  3  3   128    3727360       0.2969      4.125      5.281      1.141         3.616      4.630    
  3  3   256    1863680       0.2969      4.250      5.328      1.109         3.831      4.803    
  3  3   512     929792       0.5469      3.969      5.516      1.297         3.060      4.253    
  3  3  1024     462848       0.5156      4.188      5.453      1.297         3.229      4.205    
  3  3  2048     229376       0.5312      4.141      5.359      1.297         3.193      4.133    
  3  3  4096     114688       0.5781      4.156      5.328      1.359         3.057      3.920    
  3  3  8192      57344       0.6250      4.188      5.344      1.453         2.882      3.677    
  3  3 16384      28672       0.6406      4.094      5.328      1.453         2.817      3.667    
  3  3 32768      12288       0.5469      3.531      4.625      1.234         2.861      3.747    
  3  3 65536       4096       0.3750      2.359      3.047     0.8438         2.796      3.611    

It is unexpected that A(1:N1,1:N2,1:N3)=0.d0 is from 3 to 5 times faster! Could this be explained or is it some bug in the compiler? NOTE: This big difference happens only when first dimensions are small and the last one is big.

When it is compiled with /Od (without optimization) on the DELL Desktop the times are:

  N1 N2  N3       niter        tt1         tt2        tt3        tt4         tt2/tt4    tt3/tt4
  3  3   128    3727360        9.750      21.81      21.91      19.27         1.132      1.137    
  3  3   256    1863680        9.469      21.03      21.41      18.27         1.151      1.172    
  3  3   512     929792        9.203      20.61      21.61      18.72         1.101      1.154    
  3  3  1024     462848        9.422      21.14      21.20      18.39         1.150      1.153    
  3  3  2048     229376        9.281      20.91      21.34      18.66         1.121      1.144    
  3  3  4096     114688        9.484      20.83      21.17      18.42         1.131      1.149    
  3  3  8192      57344        9.594      21.33      21.27      18.52         1.152      1.149    
  3  3 16384      28672        9.578      21.14      21.23      18.22         1.160      1.166    
  3  3 32768      12288        8.203      18.23      17.91      15.55         1.173      1.152    
  3  3 65536       4096        5.297      12.08      12.55      10.98         1.100      1.142 

Here A(1:N1,1:N2,1:N3)=0.d0 is only 10% to 15% faster, but is persistently so (i.e. not due to random variations of the timings).

Compiled with Silverfrost FTN95 v8.80 (on the DELL Laptop) with Option /Op the timings are quite different:

  N1 N2  N3       niter        tt1         tt2        tt3        tt4         tt2/tt4    tt3/tt4
  3  3   128    3727360        2.469      2.484      10.97      13.45        0.1847     0.8153
  3  3   256    1863680        2.422      2.422      11.36      14.08        0.1720     0.8069
  3  3   512     929792        2.625      2.641      11.22      13.73        0.1923     0.8168
  3  3  1024     462848        2.563      2.594      11.33      13.55        0.1915     0.8362
  3  3  2048     229376        2.516      2.516      10.97      13.66        0.1842     0.8032
  3  3  4096     114688        2.688      2.672      11.27      13.52        0.1977     0.8335
  3  3  8192      57344        2.703      2.719      11.13      13.44        0.2023     0.8279
  3  3 16384      28672        2.672      2.734      11.27      13.44        0.2035     0.8384
  3  3 32768      12288        2.375      2.359      9.672      11.48        0.2054     0.8422
  3  3 65536       4096        1.594      1.594      6.719      7.906        0.2016     0.8498

Now A=0 is the fastest (same as the 'vector' subroutine) and A(1:N1,1:N2,1:N3) is about 5 times slower.

gFortran v12.1, with option '-O', on DELL Laptop, gives:

  N1 N2  N3       niter        tt1         tt2        tt3        tt4         tt2/tt4    tt3/tt4
  3  3   128    3727360        1.297      1.203      1.219      1.234        0.9747     0.9873    
  3  3   256    1863680        1.219      1.219      1.203      1.234        0.9873     0.9747    
  3  3   512     929792        1.344      1.359      1.328      1.391        0.9775     0.9551    
  3  3  1024     462848        1.359      1.344      1.344      1.344         1.000      1.000    
  3  3  2048     229376        1.328      1.375      1.328      1.359         1.011     0.9770    
  3  3  4096     114688        1.359      1.375      1.359      1.359         1.011      1.000    
  3  3  8192      57344        1.438      1.453      1.453      1.438         1.011      1.011    
  3  3 16384      28672        1.438      1.438      1.453      1.453        0.9892      1.000    
  3  3 32768      12288        1.234      1.234      1.344      1.250        0.9875      1.075    
  3  3 65536       4096       0.8594     0.9062     0.8750     0.8594         1.055      1.018    

The timings are practically the same for all discussed options of initialization, and they are close to Intel's best times.

18/07/2022. Did some testing and found that the way the array dimensions are defined is very important. If N1 and N2 does not have constant values (say vary in a DO loop, or 'READ') the times to initialize A by the three syntaxes are similar and much slower than the subroutine vector initialization. If N1 and N2 are assigned a constant, then we see the results presented above. If N1 and N2 are given as parameters, then Intel compiler gives all times equal to the subroutine vector initialization time. For the Silverfrost compiler - only A=0 is as fast as the 'vector' time.

0 Answers
Related