understanding pytorch quantization model output

Viewed 88

Here's two ways to use pytorch's mobilenet_v3_large.

The first not using quantization:

import torchvision
import torch

model = torchvision.models.detection.fasterrcnn_resnet50_fpn(pretrained=True)
model.eval()
x = [torch.rand(3, 300, 400), torch.rand(3, 500, 400)]
predictions = model(x)
print(predictions)

gives an output:

[{'boxes': tensor([], size=(0, 4), grad_fn=<StackBackward0>), 'labels': tensor([], dtype=torch.int64), 'scores': tensor([], grad_fn=<IndexBackward0>)}, {'boxes': tensor([], size=(0, 4), grad_fn=<StackBackward0>), 'labels': tensor([], dtype=torch.int64), 'scores': tensor([], grad_fn=<IndexBackward0>)}]

The second does use quantization:

import torch
from torchvision.io import read_image
from torchvision.models.quantization import mobilenet_v3_large, MobileNet_V3_Large_QuantizedWeights

torch.backends.quantized.engine = 'qnnpack'
img = read_image("someimage.jpg")
# Step 1: Initialize model with the best available weights
weights = MobileNet_V3_Large_QuantizedWeights.DEFAULT
model = mobilenet_v3_large(weights=weights, quantize=True)
model.eval()
# Step 2: Initialize the inference transforms
preprocess = weights.transforms()
# Step 3: Apply inference preprocessing transforms
batch = preprocess(img).unsqueeze(0)
# Step 4: Use the model and print the predicted category
prediction = model(batch)
print(prediction)

which when run, outputs something like this:

tensor([[-12.1885,  -6.0943,  -2.9548,  -4.6169,  -6.4636,  -4.0628,  -8.8644,
          -7.7563,  -6.8330,  -6.0943,  -6.2789,  -5.9096,...]])

I'm trying to figure out how to reconcile these two outputs. How can I retrieve "boxes", "labels" etc. from the quantized model?

Thanks!

torch version = 1.12.0

torchvision version = 0.13.0

1 Answers

You are importing the wrong module/weights. If you want to perform object detection, you have to import your models from the torchvision.models.detection module.

In your second code snippet, you are loading weights of only MobileNet network and using it on a MobileNet network, not an object detection network (the two are different). The output you got was basically the features/representation of your input image.

In your first code snippet, you used fasterrcnn_resnet50_fpn, which uses ResNet50 as its backbone network and not MobileNet. You cannot use MobileNet weights with a ResNet backbone network. The following code will work because the backbone network is ResNet50 and the weights are also of a ResNet50 network.

Related PyTorch documentation

import torch
from torchvision.io import read_image
from torchvision.models.quantization import ResNet50_QuantizedWeights

model = torchvision.models.detection.fasterrcnn_resnet50_fpn(weights=ResNet50_QuantizedWeights)
model.eval()
x = [torch.rand(3, 300, 400), torch.rand(3, 500, 400)]
predictions = model(x)
print(predictions)
Related