Let's have a double precision multidimensional array which needs to be initialized.
allocatable A(:,:,:)
! some code in which N1,N2,N3 are defined
allocate (A(N1,N2,N3))
We can initialize A in three syntax-wise ways:
A=0.d0
A(:,:,:)=0.0d0
A(1:N1,1:N2,1:N3)=0.0d0
Which one would give the fastest execution time? (to account for the variability of the execution time of the same code, neglect difference less than, say 20%)
Comments so far indicate that the execution time is not expected to be dependent on the syntax. However this test problem shows differently:
program MatrixInit
implicit double precision(A-H,O-Z)
allocatable A(:,:,:)
! writing file header
write(44,'(A,A)')' N1 N2 N3 niter tt1 tt2 ', &
' tt3 tt4 tt2/tt4 tt3/tt4'
! Set the small dimensions N1 and N2
! To hold, say, a stress tensor or a rotational matrix
N1=3; N2=3
! Set several examples with big N3 (say number of nodes)
do i3=7,16
N3=2**i3
allocate(A(N1,N2,N3))
! Set the number of repetitions so that the times for the examples to be similar
niter = (2**20)/(N1*N2*N3)*(2**12)
! NOTE: In the repetitions loop a duumy assignment is used so that the
! initialization is not optimized to be out of the loop
! t1 convert matrix to vector - This the fastest
call cpu_time(t0)
do iter=1,niter
A(1,1,1)=iter
call INIT(A,N1*N2*N3)
enddo
call cpu_time(t1)
! t2: A=0.0d0
do iter=1,niter
A(1,1,1)=iter
A=0.d0
enddo
call cpu_time(t2)
! t3: A(:,:)=0.0d0
do iter=1,niter
A(1,1,1)=iter
A(:,:,:)=0.d0
enddo
call cpu_time(t3)
! t4: A(1:N2,1:N3)=0.0d0
do iter=1,niter
A(1,1,1)=iter
A(1:N1,1:N2,1:N3)=0.d0
enddo
call cpu_time(t4)
! Times
tt1=t1-t0 ! Time for vector initialization
tt2=t2-t1 ! Time for A=0.d0
tt3=t3-t2 ! Time for A(:,:,:)=0.d0
tt4=t4-t3 ! Time for A(1:N1,1:N2,1:N3)
write(44,9) N1,N2,N3, niter, tt1,tt2,tt3,tt4, tt2/tt4,tt3/tt4
deallocate (A)
enddo
9 format(2i3,i6, 2x,i9, 6x, 4G11.4, 3x, 2G11.4)
end
subroutine INIT(A,N)
implicit double precision(A-H,O-Z)
dimension A(N)
A=0.0D0
return
end
Compilation is done with Intel Fortran with Option /O2. Below is the output from the WRITE(44,...) when the program is compiled and run on two computers:
On DELL Laptop: Intel(R) Core(TM) i7-3740QM CPU @ 2.70GHz 2.70 GHz, RAM 16.0GB
Intel(R) Visual Fortran Intel(R) 64 Compiler for applications running on IA-32, Version 18.0.0.124
N1 N2 N3 niter tt1 tt2 tt3 tt4 tt2/tt4 tt3/tt4
3 3 128 3727360 0.6406 6.031 7.297 1.406 4.289 5.189
3 3 256 1863680 0.7188 6.266 7.219 1.234 5.076 5.848
3 3 512 929792 0.8906 5.984 7.172 1.328 4.506 5.400
3 3 1024 462848 0.9219 6.000 7.250 1.359 4.414 5.333
3 3 2048 229376 0.9531 5.875 7.109 1.312 4.476 5.417
3 3 4096 114688 0.9688 5.953 7.047 1.391 4.281 5.067
3 3 8192 57344 1.297 6.109 7.172 1.500 4.073 4.781
3 3 16384 28672 1.312 5.906 7.422 1.438 4.109 5.163
3 3 32768 12288 1.125 5.047 6.062 1.266 3.988 4.790
3 3 65536 4096 0.7656 3.453 4.047 0.8906 3.877 4.544
DELL Desktop: Intel(R) Core(TM) i7-6700 CPU @ 3.40GHz 3.41 GHz, RAM 32.0GB
Intel(R) Visual Fortran Intel(R) 64 Compiler for applications running on Intel(R) 64, Version 19.1.3.311
N1 N2 N3 niter tt1 tt2 tt3 tt4 tt2/tt4 tt3/tt4
3 3 128 3727360 0.2969 4.125 5.281 1.141 3.616 4.630
3 3 256 1863680 0.2969 4.250 5.328 1.109 3.831 4.803
3 3 512 929792 0.5469 3.969 5.516 1.297 3.060 4.253
3 3 1024 462848 0.5156 4.188 5.453 1.297 3.229 4.205
3 3 2048 229376 0.5312 4.141 5.359 1.297 3.193 4.133
3 3 4096 114688 0.5781 4.156 5.328 1.359 3.057 3.920
3 3 8192 57344 0.6250 4.188 5.344 1.453 2.882 3.677
3 3 16384 28672 0.6406 4.094 5.328 1.453 2.817 3.667
3 3 32768 12288 0.5469 3.531 4.625 1.234 2.861 3.747
3 3 65536 4096 0.3750 2.359 3.047 0.8438 2.796 3.611
It is unexpected that A(1:N1,1:N2,1:N3)=0.d0 is from 3 to 5 times faster! Could this be explained or is it some bug in the compiler? NOTE: This big difference happens only when first dimensions are small and the last one is big.
When it is compiled with /Od (without optimization) on the DELL Desktop the times are:
N1 N2 N3 niter tt1 tt2 tt3 tt4 tt2/tt4 tt3/tt4
3 3 128 3727360 9.750 21.81 21.91 19.27 1.132 1.137
3 3 256 1863680 9.469 21.03 21.41 18.27 1.151 1.172
3 3 512 929792 9.203 20.61 21.61 18.72 1.101 1.154
3 3 1024 462848 9.422 21.14 21.20 18.39 1.150 1.153
3 3 2048 229376 9.281 20.91 21.34 18.66 1.121 1.144
3 3 4096 114688 9.484 20.83 21.17 18.42 1.131 1.149
3 3 8192 57344 9.594 21.33 21.27 18.52 1.152 1.149
3 3 16384 28672 9.578 21.14 21.23 18.22 1.160 1.166
3 3 32768 12288 8.203 18.23 17.91 15.55 1.173 1.152
3 3 65536 4096 5.297 12.08 12.55 10.98 1.100 1.142
Here A(1:N1,1:N2,1:N3)=0.d0 is only 10% to 15% faster, but is persistently so (i.e. not due to random variations of the timings).
Compiled with Silverfrost FTN95 v8.80 (on the DELL Laptop) with Option /Op the timings are quite different:
N1 N2 N3 niter tt1 tt2 tt3 tt4 tt2/tt4 tt3/tt4
3 3 128 3727360 2.469 2.484 10.97 13.45 0.1847 0.8153
3 3 256 1863680 2.422 2.422 11.36 14.08 0.1720 0.8069
3 3 512 929792 2.625 2.641 11.22 13.73 0.1923 0.8168
3 3 1024 462848 2.563 2.594 11.33 13.55 0.1915 0.8362
3 3 2048 229376 2.516 2.516 10.97 13.66 0.1842 0.8032
3 3 4096 114688 2.688 2.672 11.27 13.52 0.1977 0.8335
3 3 8192 57344 2.703 2.719 11.13 13.44 0.2023 0.8279
3 3 16384 28672 2.672 2.734 11.27 13.44 0.2035 0.8384
3 3 32768 12288 2.375 2.359 9.672 11.48 0.2054 0.8422
3 3 65536 4096 1.594 1.594 6.719 7.906 0.2016 0.8498
Now A=0 is the fastest (same as the 'vector' subroutine) and A(1:N1,1:N2,1:N3) is about 5 times slower.
gFortran v12.1, with option '-O', on DELL Laptop, gives:
N1 N2 N3 niter tt1 tt2 tt3 tt4 tt2/tt4 tt3/tt4
3 3 128 3727360 1.297 1.203 1.219 1.234 0.9747 0.9873
3 3 256 1863680 1.219 1.219 1.203 1.234 0.9873 0.9747
3 3 512 929792 1.344 1.359 1.328 1.391 0.9775 0.9551
3 3 1024 462848 1.359 1.344 1.344 1.344 1.000 1.000
3 3 2048 229376 1.328 1.375 1.328 1.359 1.011 0.9770
3 3 4096 114688 1.359 1.375 1.359 1.359 1.011 1.000
3 3 8192 57344 1.438 1.453 1.453 1.438 1.011 1.011
3 3 16384 28672 1.438 1.438 1.453 1.453 0.9892 1.000
3 3 32768 12288 1.234 1.234 1.344 1.250 0.9875 1.075
3 3 65536 4096 0.8594 0.9062 0.8750 0.8594 1.055 1.018
The timings are practically the same for all discussed options of initialization, and they are close to Intel's best times.
18/07/2022. Did some testing and found that the way the array dimensions are defined is very important. If N1 and N2 does not have constant values (say vary in a DO loop, or 'READ') the times to initialize A by the three syntaxes are similar and much slower than the subroutine vector initialization. If N1 and N2 are assigned a constant, then we see the results presented above. If N1 and N2 are given as parameters, then Intel compiler gives all times equal to the subroutine vector initialization time. For the Silverfrost compiler - only A=0 is as fast as the 'vector' time.