I'm trying to do matrix-vector multiplication in parallel in C++ using std::for_each and std::transform_reduce. I can get the correct answers, however my runtime doesn't change when using std::par vs std::seq for the execution policy.
For example,
std::vector<double> vectorToMultiplyBy(10);
std::vector<std::vector<double>> someMatrix(10, 10);
for (auto it = someMatrix.begin(); it != someMatrix.end(); ++it)
{
auto i = std::transform_reduce(std::execution::par, it->begin(), it->end(), vectortoMulitplyBy.begin());
}
This code provides the same results but the same or slower runtime when std::execution::par is exchanged with std::execution::seq.
Same idea with for_each:
std::vector<double> vectorToMultiplyBy(10);
std::vector<std::vector<double>> matrix(10, 10);
std::for_each(std::execution::par, someMatrix.begin(), someMatrix.end(), [](std::vector<double> currentRow)
{
std::transform_reduce(std::execution::par, currentRow.begin(), currentRow.end(), vectorToMultiplyBy.begin());
});
Switching std::execution::par with std::execution::seq in the above will not change anything if not make things slower. I have tried with up to 32768x32768 matrix (and a 32768x1 vector in that case)
I have confirmed that I am using c++17. I am using g++11 on Apple M1 processor (10 cores).
Timing was done with std:chrono::high_resolution_clock.