Loop dependencis when parallelizing loops with OpenMP

Viewed 63

I tried parallelizing the following loop:

    #pragma omp parallel for shared(v, g, p)
    for (i=1; i<=imax-1; i++) {
        for (j=1; j<=jmax-1; j++) {  // combined loops
            /* only if both adjacent cells are fluid cells */
            if ((flag[i][j] & C_F) && (flag[i][j+1] & C_F)) {
                v[i][j] = g[i][j]-(p[i][j+1]-p[i][j])*del_t/dely;
            }
            if ((flag[i][j] & C_F) && (flag[i+1][j] & C_F)) {
                u[i][j] = f[i][j]-(p[i+1][j]-p[i][j])*del_t/delx;
            }
        }
    }

But my program does not run as expected, probably because of loop dependencies. Is there a way to parallelize this loop with reductions? Any help will be greatly appreciated.

1 Answers

But my program does not run as expected, probably because of loop dependencies. Is there a way to parallelize this loop with reductions?

There is a race-condition in one of the shared variables (i.e., the loop iterator j) being updated concurrently by multiple threads.

   #pragma omp parallel for shared(v, g, p)
   for (i=1; i<=imax-1; i++) {
        for (j=1; j<=jmax-1; j++) {  // <---- Race-condition 
            /* only if both adjacent cells are fluid cells */
            if ((flag[i][j] & C_F) && (flag[i][j+1] & C_F)) {
                v[i][j] = g[i][j]-(p[i][j+1]-p[i][j])*del_t/dely;
            }
            if ((flag[i][j] & C_F) && (flag[i+1][j] & C_F)) {
                u[i][j] = f[i][j]-(p[i+1][j]-p[i][j])*del_t/delx;
            }
        }
    }

The variable i of the outermost loop does not have a race-condition because it belongs to the loop being parallelized, and in that case, the OpenMP standard states that that variable will be implicitly private.

To solve the race-condition you need to make the variable j of the innermost loop private. Either by doing:

   #pragma omp parallel for shared(v, g, p) private(j)
   for (i=1; i<=imax-1; i++) {
        for (j=1; j<=jmax-1; j++) {

or (depending in your compiler version):

   #pragma omp parallel for shared(v, g, p)
   for (i=1; i<=imax-1; i++) {
        for (int j=1; j<=jmax-1; j++) {
Related