While I've read PyTorch Quantization tutorial, I ran into 'torch.quantization.fuse_modules'.
The documentation said
Next, we'll "fuse modules"; this can both make the model faster by saving on memory access while also improving numerical accuracy.
I don't understand why 'fuse modules' can make the model faster and improve numerical accuracy.
Would you elaborate this?