I put this code only as an example so that you can understand what I am looking for:
double *f = malloc(sizeof(double) * nx * ny);
double *f2 = malloc(sizeof(double) * nx * ny);
for ( i = process * (nx/totalProcesses); i < (process + 1) * (nx/totalProcesses); i++ )
{
for ( j = 0; j < ny; j++ )
{
f2[i*ny + j] = j*i;
}
}
MPI_Allreduce( f2, f, nx*ny, MPI_DOUBLE, MPI_SUM, MPI_COMM);
And yes, it works, in the end I have the correct result in 'f' and that is what I want, but I would like to know if there is a better or more direct way to achieve the same in order to get efficiency. I tried it with allgather but couldn't get correct result.