Doing a zip transform with a c++ SIMD header library we might have the following sudo code.
// using xsimd
binary_op = [](const auto& a, const auto& b){ return ...; }
float* a, b, res;
...
for(auto i = 0; i < ...; i += batch_size)
{
auto batch_a = xs::load_aligned(a += i);
auto batch_b = xs::load_aligned(b += i);
auto batch_res = binary_op(batch_a, batch_b);
batch_res.store_aligned(res += i);
}
I was wondering whether adjacent transform
a0
a1 a0 -> a1+a0
a2 a1 -> a2+a1
a3 a2 -> a3+a2
a3
could be speed up since we might be able to call load_aligned only once per iteration.
auto batch_a1 = xs::load_aligned(...);
...
for (...)
{
auto batch_a0 = batch_a1;
batch_a1 = xs::load_aligned(...);
auto batch_b0 = ...; // somehow create from batch_a0 and batch_a1
auto batch_res = binary_op(batch_a0, batch_b0);
batch_res.store_aligned(...);
}
Perhaps someone could suggest how to perform the following kind of operation in simd intrinsics:
([a0, a1, a2, a3], [a4, a5, a6, a7]) -> [a1, a2, a3, a4]
And would this even be likely to cause a speed up?