I could not spot a problem with my use of std::reduce() function from the <numeric> STL header.
Since I have found workaround, I will show first expected behavior:
uint64_t f(uint64_t n)
{
return 1ull;
}
uint64_t solution(uint64_t N) // here N == 10000000
{
uint64_t r(0);
// persistent array of primes
const auto& app = YTL::AccumulativePrimes::global().items();
auto citEnd = std::upper_bound(app.cbegin(), app.cend(), 2*N);
auto citBegin = std::lower_bound(app.cbegin(), citEnd, N);
std::vector<uint64_t> v(citBegin, citEnd);
std::for_each(std::execution::par,
v.begin(), v.end(),
[](auto& p)->void {p = f(p); });
r = std::reduce(std::execution::par, v.cbegin(), v.cend(), 0);
return r; // here is correct answer: 606028
}
However, if I want to avoid intermediate vector and instead apply binary operator on the spot in reduce() itself, also in parallel, it gives me a different answer each time:
uint64_t f(uint64_t n)
{
return 1ull;
}
uint64_t solution(uint64_t N) // here N == 10000000
{
uint64_t r(0);
// persistent array of primes
const auto& app = YTL::AccumulativePrimes::global().items();
auto citEnd = std::upper_bound(app.cbegin(), app.cend(), 2*N);
auto citBegin = std::lower_bound(app.cbegin(), citEnd, N);
// bug in parallel reduce?!
r = std::reduce(std::execution::par,
citBegin, citEnd, 0ull,
[](const uint64_t& r, const uint64_t& v)->uint64_t { return r + f(v); });
return r; // here the value of r is different every time I run!!
}
Could anyone explain why the latter usage is wrong?
I am using MS C++ compiler cl.exe: Version 19.28.29333.0;
Windows SDK version: 10.0.18362.0;
Platform Toolset: Visual Studio 2019 (v142)
C++ language standard: Preview - Features from the Latest C++ Working Draft (/std:c++latest)
Computer: Dell XPS 9570 i7-8750H CPU @ 2.20GHz, 16GB RAM OS: Windows 10 64bit