I'm trying to do matrix multiplication in parallel using the NVIDIA HPC SDK's stdpar implementation, and ran into a problem.
Is there any way I can accomplish the following without having to capture the variables by reference inside the lambdas? My goal is to run the loops on the GPU as well.
I'm trying to compile this using the nvc++ compiler using the -stdpar flag, which does not allow capture by reference, as it would likely cause an illegal memory access when run on the GPU.
std::vector<std::vector<T>> result;
std::for_each(std::execution::par_unseq, A.begin(), A.end(),
[&](auto a) {
std::vector<T> tmp(A.size());
tmp.reserve(A.size());
std::for_each(std::execution::par_unseq, tB.begin(), tB.end(),
[&](auto b) {
tmp.push_back(std::transform_reduce(
std::execution::par_unseq,
a.begin(), a.end(), b.begin(), 0.0)
);
});
result.push_back(tmp);
});