Does anyone know if there's any way to reduce the amount of gc and/or dynamic dispatch for a gradient call in flux? I've tried using FastClosures.jl, as well as wrapping loss into a callable struct to prevent closures and consequently Core.Box calls, but nothing seems to make an appreciable difference.
MWE:
using Flux
function get_grad(dnn, x, y)
p = params(dnn)
g = gradient(p) do
Flux.mse(dnn(x),y)
end
return g
end
const DNN = Dense(3, 2)
const X = rand(Float32, 3)
const Y = rand(Float32, 2)
@profiler for _ in 1:10_000; get_grad(DNN, X, Y); end
Profile (orange - dynamic dispatch, red - garbage collection, dynamic dispatch)