Why does 'torch.quantization.fuse_modules' save on memory access and improve numerical accuracy?

Viewed 10

While I've read PyTorch Quantization tutorial, I ran into 'torch.quantization.fuse_modules'.

The documentation said

Next, we'll "fuse modules"; this can both make the model faster by saving on memory access while also improving numerical accuracy.

I don't understand why 'fuse modules' can make the model faster and improve numerical accuracy.

Would you elaborate this?

0 Answers
Related