I am trying to test the first function on the github page of SIMD.jl (https://github.com/eschnett/SIMD.jl). This is the script:
using SIMD
using BenchmarkTools
function vadd!(xs::Vector{T}, ys::Vector{T}, ::Type{Vec{N,T}}) where {N, T}
@assert length(ys) == length(xs)
@assert length(xs) % N == 0
lane = VecRange{N}(0)
@inbounds for i in 1:N:length(xs)
xs[lane + i] += ys[lane + i]
end
end
b1 = Vector{Float64}([2.0,2.0,3.0,4.0,2.0,2.0,3.0,4.0])
b2 = Vector{Float64}([2.0,2.0,3.0,4.0,2.0,2.0,3.0,4.0])
c1 = Vec{8,Float64}((2.0,2.0,3.0,4.0,2.0,2.0,3.0,4.0))
c1t = typeof(c1)
@btime b1+b2
println(b1+b2)
@btime vadd!($b1, $b2, Vec{8, Float64})
println(b1)
@btime vadd!($b1, $b2, $c1t)
The results I get are not fully clear to me:
57.739 ns (1 allocation: 144 bytes)
[4.0, 4.0, 6.0, 8.0, 4.0, 4.0, 6.0, 8.0]
18.400 ns (0 allocations: 0 bytes)
[2.1001006e7, 2.1001006e7, 3.1501509e7, 4.2002012e7, 2.1001006e7, 2.1001006e7, 3.1501509e7, 4.2002012e7]
127.009 ns (0 allocations: 0 bytes)
The first call takes 57.739 ns, which is reasonable since the '+' operation should be already vectorised and it takes some extra time to allocate the memory required to allocate the output vector.
The second call seems successful in terms of computational time (18.400 ns). However, the output stored in b1 is totally wrong. Any idea?
The third is unclear as well to me. I am simply passing the type Vec using an auxiliary variables and the resulting function is almost one order of magnitude slower. The output result is wrong as well, but this is in line with the result of the second call. Any clue on this speed loss?
Thank you in advance.