Is there a way to reduce garbage collection and/or dynamic dispatch for Julia Flux gradient calls?

Viewed 116

Does anyone know if there's any way to reduce the amount of gc and/or dynamic dispatch for a gradient call in flux? I've tried using FastClosures.jl, as well as wrapping loss into a callable struct to prevent closures and consequently Core.Box calls, but nothing seems to make an appreciable difference.

MWE:

using Flux

function get_grad(dnn, x, y)
    p = params(dnn)
    g = gradient(p) do
        Flux.mse(dnn(x),y)
    end
    return g
end

const DNN = Dense(3, 2)
const X = rand(Float32, 3)
const Y = rand(Float32, 2)

@profiler for _ in 1:10_000; get_grad(DNN, X, Y); end

Profile (orange - dynamic dispatch, red - garbage collection, dynamic dispatch)

Zygote\src\compiler\interface.jl

0 Answers
Related