Let's assume the following input data
struct Bar {
std::vector<size_t> i1s;
std::vector<size_t> i2s;
std::vector<size_t> i3s;
std::vector<size_t> i4s;
std::vector<size_t> i5s;
};
struct Foo_tuple {
size_t i1;
size_t i2;
size_t i3;
size_t i4;
size_t i5;
Foo_tuple() = default;
Foo_tuple(size_t i1 ,size_t i2 ,size_t i3 ,size_t i4 ,size_t i5) : i1(i1) ,i2(i2) ,i3(i3) ,i4(i4) ,i5(i5) {}
};
struct Foo {
std::vector<Foo_tuple> tuples;
void to_bar(Bar *bar);
void to_foo(Foo *foo);
};
with the following implementations
void Foo::to_foo(Foo *foo) {
size_t tuplesToProcess = tuples.size();
auto inputBegin = tuples.begin();
auto outputBegin = foo->tuples.begin();
std::copy_n(inputBegin, tuplesToProcess, outputBegin);
};
void Foo::to_bar(Bar *bar) {
size_t tuplesToProcess = tuples.size();
for (size_t tuple = 0; tuple < tuplesToProcess; ++tuple) {
bar->i1s[tuple] = tuples[tuple].i1;
bar->i2s[tuple] = tuples[tuple].i2;
bar->i3s[tuple] = tuples[tuple].i3;
bar->i4s[tuple] = tuples[tuple].i4;
bar->i5s[tuple] = tuples[tuple].i5;
}
};
Now, I observe that both to_foo() and to_bar() calls in
Foo inputTuples;
Foo outputTuples;
// do proper population of the inputTuple's and resizing of the outputTuple's here
inputTuples.to_foo(&outputTuples);
as well as
Foo inputTuples;
Bar outputTuples;
// do proper population of the inputTuple's and resizing of the outputTuple's here
inputTuples.to_bar(&outputTuples);
lead to exactly the same performance (measured with google benchmark).
This is somehow surprising to me, since the foo to foo transformation can be performed by a sequential copy of memory (I also confirmed that it results in a memcpy) while the bar to foo transformation has to read one tuple from one location and then has to write five different locations, which should result in way more random memory access.
Additionally I observed, that this behavior only holds as long as
The number of fields in the struct is less or equal than 8. With more fields, the foo to bar transformation starts to get significantly slower the more fields I add. This plot has a fixed tuple count of 8'192 and varies the field count

The struct has at least 16 tuples. When having less than 16 tuples, the foo to bar transformation starts to get significantly slower the less tuples are in the input. This plot has a fixed field count of 20 and varies the tuple count

Can anyone explain to me where these three effects come from?
Edit:
The code was compiled with gcc 12.0.1 and -O3 on a 64 Bit platofrm.