std::execution::par execution policy not providing speedup with std::for_each and std::transform_reduce

Viewed 101

I'm trying to do matrix-vector multiplication in parallel in C++ using std::for_each and std::transform_reduce. I can get the correct answers, however my runtime doesn't change when using std::par vs std::seq for the execution policy.

For example,

std::vector<double> vectorToMultiplyBy(10);
std::vector<std::vector<double>> someMatrix(10, 10);

for (auto it = someMatrix.begin(); it != someMatrix.end(); ++it)
{
   auto i = std::transform_reduce(std::execution::par, it->begin(), it->end(), vectortoMulitplyBy.begin());
}

This code provides the same results but the same or slower runtime when std::execution::par is exchanged with std::execution::seq.

Same idea with for_each:

std::vector<double> vectorToMultiplyBy(10);
std::vector<std::vector<double>> matrix(10, 10);

std::for_each(std::execution::par, someMatrix.begin(), someMatrix.end(), [](std::vector<double> currentRow)
{

   std::transform_reduce(std::execution::par, currentRow.begin(), currentRow.end(), vectorToMultiplyBy.begin());

});

Switching std::execution::par with std::execution::seq in the above will not change anything if not make things slower. I have tried with up to 32768x32768 matrix (and a 32768x1 vector in that case)

I have confirmed that I am using c++17. I am using g++11 on Apple M1 processor (10 cores).

Timing was done with std:chrono::high_resolution_clock.

0 Answers
Related