This question is based on this video on YouTube made with the purpose of reviewing this project.
In the video, the host is analyzing the project and found out that the following block of code is a cause of performance issues:
std::optional<HitRecord> HittableObjectList::Hit(const Ray &r, float t_min, float t_max) const {
float closest_so_far = t_max;
return std::accumulate(begin(objects), end(objects), std::optional<HitRecord>{},
[&](const auto &temp_value, const auto &object) {
if(auto temp_hit = object -> Hit(r, t_min, closest_so_far); temp_hit) {
closest_so_far = temp_hit.value().t;
return temp_hit;
}
return temp_value;
});
}
I would assume that the std::accumulate function would function similarly to a for loop. Unhappy with the performance hit there (and because, for some reason, the profiler wouldn't profile the lambda code[a limitation, perhaps?]), the reviewer changed the code to this:
std::optional<HitRecord> HittableObjectList::Hit(const Ray &r, float t_min, float t_max) const {
float closest_so_far = t_max;
std::optional<HitRecord> record{};
for(size_t i = 0; i < objects.size(); i++) {
const std::shared_ptr<HittableObject> &object = objects[i];
if(auto temp_hit = object -> Hit(r, t_min, closest_so_far); temp_hit) {
closest_so_far = temp_hit.value().t;
record = temp_hit;
}
}
return record;
}
With this change the time to completion went from 7 minutes and 30 seconds to 22 seconds.
My questions are:
- Aren't both blocks of code identical? Why does
std::accumulategive such enormous penalty here? - Would the performance be better if instead of using
autos, using the explicit type?
The reviewer did mention suggestions such as avoiding the use of std::optionals and std:shared_ptrs here due to the amount of calls made and to execute this code on the GPU instead, but for now I'm only interested in those points mentioned earlier.