Is there a way to *not* have to capture variables by reference in this?

Viewed 102

I'm trying to do matrix multiplication in parallel using the NVIDIA HPC SDK's stdpar implementation, and ran into a problem.

Is there any way I can accomplish the following without having to capture the variables by reference inside the lambdas? My goal is to run the loops on the GPU as well.

I'm trying to compile this using the nvc++ compiler using the -stdpar flag, which does not allow capture by reference, as it would likely cause an illegal memory access when run on the GPU.

std::vector<std::vector<T>> result;
std::for_each(std::execution::par_unseq, A.begin(), A.end(),
                  [&](auto a) {
                      std::vector<T> tmp(A.size());
                      tmp.reserve(A.size());
                      std::for_each(std::execution::par_unseq, tB.begin(), tB.end(),
                                    [&](auto b) {
                                        tmp.push_back(std::transform_reduce(
                                            std::execution::par_unseq, 
                                            a.begin(), a.end(), b.begin(), 0.0)
                                        );
                                    });
                      result.push_back(tmp);
                  });
1 Answers

I have a similar question. I don't have enough reputation to comment, but per the NVIDIA docs:

For example, std::vector uses dynamically allocated memory, which is accessible from the GPU when using stdpar. Iterating over the contents of std::vector in a C++ Parallel Algorithm works as expected:

It does say in the docs that you can't do capture by reference, but they were talking about an std::array in that context which is not dynamically allocated internally.

So I guess my point is, if you were using std::vector which is dynamically allocated internally, (and you are) it might work according to the docs. Did you try it?

As another side note, even if it did not have a memory access issue, I don't think it would be a good idea to push_back within a parallel loop because it would be a race condition, which means that the result of that vector you're pushing into would depend on which thread ran at what time. It may have correct answers but they might be out of order.

I'm not sure how to avoid the race condition, I'm trying to figure that exact thing out myself as well with code similar to yours but I'm not using NVIDIA HPC.

I understand this may not fully answer your question but I'm not able to just comment due to reputation. I hope you figured out a solution.

Related