I need to speed up a multi-dimensional array assignment with matrix multiplication over some indices. The original cycle is
FORALL (ia=1:2,ib=1:2,ic=1:2) U(:,ia,:,ib,ic)=
+ W(ic,ia,ib,1)*matmul(FS(:,ia,:,1),transpose(F(:,ib,:,1)))+
+ W(ic,ia,ib,2)*matmul(FS(:,ia,:,2),transpose(F(:,ib,:,2)))
The matrix multiplication is over the large dimension >100. What is the best way to increase speed using OMP directives on a workstation with parallel Intel Xeon cores ? Say,
!$OMP DO
DO ic=1,2
DO ib=1,2
DO ia=1,2
U(:,ia,:,ib,ic)=
+ W(ic,ia,ib,1)*matmul(FS(:,ia,:,1),transpose(F(:,ib,:,1)))+
+ W(ic,ia,ib,2)*matmul(FS(:,ia,:,2),transpose(F(:,ib,:,2)))
ENDDO
ENDDO
ENDDO
!$OMP END DO
Will this work or there are better alternatives ? Thank you