I want to use OpenMP to parallelize a for-loop calculator which does something like:
B = (int*)malloc(sizeof(int) * N); //N is known
for(i=0;i<500000;i++)
{
for(j=0;j<M;j++) B[j]=i+j; //M is different from N, but M <= N;
some operations on B which produce a variable L;
printf("%d\n",L);
}
I don't need to re-allocate B as its values will be defined for each iteration accordingly. The operations will only use B[0] to B[M-1]. This saves a lot of time in allocating and initialization of B.
In order to use openmp, I changed the code to this:
#pragma omp parallel num_threads(32) private(i,j,B,M,L)
{
B = (int*)malloc(sizeof(int) * N); //N is known
#pragma omp parallel for
for(i=0;i<500000;i++)
{
for(j=0;j<M;j++) B[j]=i+j; //M is different from N, but M <= N;
some operations on B which produce a variable L;
printf("%d\n",L);
}
}
It runs really slow compared to the first one, as it creates a new B array for each thread (so 500000 times). Is there a way to avoid this using openmp?