How to add a single piece of information to a UNet imput

Viewed 47

I am performing segmentation using segmentation_models, which is a wrapper for keras. This is the blurb that defines my UNet:

jaccard_loss = sm.losses.JaccardLoss(class_weights=class_weights)
focal_loss = sm.losses.CategoricalFocalLoss()
total_loss = jaccard_loss + (1 * focal_loss)
metrics = [sm.metrics.IOUScore()]    
model = sm.Unet(BACKBONE1, encoder_weights=None,classes=n_classes, activation='softmax',input_shape=(None, None, num_channels))
model.compile(opt, total_loss, metrics=metrics)

My question is relatively simple, I'm feeding in a stack of slices into the UNet, but there's a lot of spatial information that's missing (i.e., just the physical location of the slice). I would like to feed this into the model to see if this helps improve segmentation. The easiest thing to do would be to just have another channel which has an image that is all the same value (i.e., a uniform image of 0 to 1 depending on physical location). I have a feeling this is not the best way though, so I was wondering if anyone had any good ideas or has done something similar before? Thank you very much in advance for your help.

1 Answers

Relative location can be a very useful cue for segmentation. Adding it as an additional channel can be very beneficial.

For instance, in our recent work:
O. Frank et al., Integrating Domain Knowledge into Deep Networks for Lung Ultrasound with Applications to COVID-19 in IEEE Transactions on Medical Imaging (2021).
We augmented a Lung Ultrasound (LUS) frame with an additional relative position channel that proved very useful for the classification and segmentation of COVID-19 related bio-markers.
This work exemplifies how using additional channels for introducing ``domain knowledge" can be very efficient and beneficial.

Another related work, on the analysis of COVID-19 in chest Xray (CXR):
Keidar, D., Yaron, D., Goldstein, E. et al. COVID-19 classification of X-ray images using deep neural networks Eur Radiol (2021).
Also augmented the raw XCR frame with positional information relative to the location of the lungs.

You can think of better ways of encoding position than a 0/1 mask, like positional encoding/embeddings used in vision transformers (ViT) architecture.

Related