I am currently working on the implementation of a multidimensional array iterator. Considering the iteration over two contiguous ranges (for std::equal, std::copy purposes) that represent compatible data with different alignements (row vs col major in 2D), I would like to find the stride order for each iterator giving the fastest execution time.
For example:
row of vector components = A -> m elements
row of vectors = B -> n elements
2D plan of vectors = C -> 3 elements
row of plan of vectors = D -> 10 elements
given the datas ordered by ascending strides:
first array: B | A | C | D
second array: B | A | D | C
Obviously, we can iterate over both iterators by bunches of m*n elements. Then:
If we choose the first array convention, the first iterator is contiguous and the second one
will perform (3 - 1)*(10 - 1) jumps forward with a stride of 10 and (10 - 1) jumps backward.
If we choose the second array convention, the second iterator is contiguous and the first one will
perform (10 - 1)*(3 - 1) jumps forward with a stride of 3 and (3 - 1) jumps backward.
=> The second convention is better at everything in this example.
Since I have to take a lot of factors to take into account like the memory back and forth, the contiguity and the iterator implementation itself (which is not trivial), I want to perform an experimental plan. But I also know everything at compile time (the sizes and the strides) so it would be cool to perform the experimental plan at compile time for each template instantiation. My question is:
Is it possible to evaluate the runtime cost of some instructions at compile time, when everything is known at compile time except the memory address of the input array ?