I'm trying to improve a Fortran 77 code by vectorizing a for loop. I'm fairly new to vectorization and while I can get the code to vectorize, an optimization report tells me that some of my arrays have unaligned access. This makes, as far as I understand it, the vectorization less efficient. I have manually added padding to my arrays in order to align the data and this seems to work for most of my arrays, but not for all of them (see example code).
SUBROUTINE SOURCE
PARAMETER(NLINES=8)
PARAMETER(NLM=NLINES-1)
PARAMETER(NZ=18)
COMMON/TEST/h(-nlines:nlines+15,-nlines:nlines)
&,hr(-nlines:nlines+15,-nlines:nlines)
COMMON/PRO/P2(-nlm:nlm+1,-nlm:nlm)
dimension SNT(-nlm:nlm+1,-nlm:nlm,0:NZ+1)
DO iy=-nlm,nlm
DO ix=-nlm,nlm
H(ix,iy)=SNT(ix,iy,1)
& +SNT(ix,iy,0)
& +SNT(ix,iy,2)
HR(ix,iy)=P2(ix,iy)*H(ix,iy)
enddo
enddo
END
And the relevant part of the optimization report:
LOOP BEGIN at SOURCE.FPP(13,8)
remark #15389: vectorization support: reference h(ix,iy) has unaligned access [ SOURCE.FPP(14,9) ]
remark #15388: vectorization support: reference snt(ix,iy,1) has aligned access [ SOURCE.FPP(14,9) ]
remark #15388: vectorization support: reference snt(ix,iy,0) has aligned access [ SOURCE.FPP(14,9) ]
remark #15388: vectorization support: reference snt(ix,iy,2) has aligned access [ SOURCE.FPP(15,9) ]
remark #15389: vectorization support: reference hr(ix,iy) has unaligned access [ SOURCE.FPP(18,9) ]
remark #15388: vectorization support: reference p2(ix,iy) has aligned access [ SOURCE.FPP(18,9) ]
remark #15389: vectorization support: reference h(ix,iy) has unaligned access [ SOURCE.FPP(18,9) ]
remark #15381: vectorization support: unaligned access used inside loop body
remark #15305: vectorization support: vector length 4
remark #15399: vectorization support: unroll factor set to 3
remark #15300: LOOP WAS VECTORIZED
remark #15448: unmasked aligned unit stride loads: 4
remark #15450: unmasked unaligned unit stride loads: 1
remark #15451: unmasked unaligned unit stride stores: 2
remark #15475: --- begin vector cost summary ---
remark #15476: scalar cost: 15
remark #15477: vector cost: 4.250
remark #15478: estimated potential speedup: 2.340
remark #15488: --- end vector cost summary ---
remark #25456: Number of Array Refs Scalar Replaced In Loop: 3
remark #25015: Estimate of max trip count of loop=1
LOOP END
Why don't h(ix,iy) and hr(ix,iy) have aligned access? And how can I align them for better vectorization?
Probably also relevant information:
Intel(R) Fortran Intel(R) 64 Compiler for applications running on Intel(R) 64, Version 19.0.0.117 Build 20180804
Compiler options: -qopt-report=5 -c -DLINUX -O3 -w -assume nounderscore -xCORE-AVX2 -o SOURCE.o193
Thanks in advance for your time and help!