Semantic segmentation without labels in a single class

Viewed 1167

I am kind of new to semantic segmentation. I am trying to perform segmentation of images having defects. I have the defect images annotated using a annotation tool and I created the mask for each image. I wanted to predict If an image has defect and where exactly it is located. But my problem is my defects does not look same in all the images. Example: Defects on steel- Steel breakage, erroded surface etc. I am just trying to classify if the image has defect or not and where it is located. So is it wrong to train the neural network with these all types considered as defects even though not everything lookalike?

I thought to do a binary segmentation of defect to no defect. If I am not correct how can I perform segmentation for defect and non defect images?

2 Answers

You first have to well define your problem and your objectives:

  • If you only want to detect if your image has a defect or not, it's a binary classification problem and you affect a label (0 or 1) to each image.
  • If you want to localise the defect approximatively (like a bounding box), it's an object detection problem and it can be realised with one or more classes.
  • If you want to localise precisely the defect (in order to performe measures for instance) the best is semantic segmentation or instance segmentation.
  • If you want to classify the defect, you will need to create classes for each defect you want to classify.

There is no magical solution because it depends of the objectives of your project. I can give you the following advices because I made an internship on a similar project :

  • Look carefully at your data, if you have thousands of images it will take a long to create your semantic segmentation dataset. Be smarter by using data augmentation techniques.
  • If you want to classify the defects, be sure to have enough defects of each type to train your network. If your network only sees one defect type per epoch, it can't learn to detect it.
  • Be sure that your network can detect the defects you're providing (not a scratch of two pixels for instance or alignement defects).

Performing semantic segmentation to only knows if there is a defect or not seems overkill because it's a long and complex process (rebuilding the image, memory of intermediaries images in Unet, lot of computations). If you really want to apply this method, you may create a threshold to detect if the number of detected pixels as defect allows to classify the image as 'presenting a defect' or not.

One class should be enough for your use-case. If you want to be able to distinguish between different types of defects though, you could try creating attributes for that class. So the class would be if a pixel has a defect or not, and the attribute would be breakage, eroded pixel, etc. Then you could train a model to detect a crack on the semantic class and another one to identify which type of defect it is.

Make sure to use an annotation tool that supports creating attributes. Personally, I use hasty.ai as their automation assistants are great! But I guess most tools should be able to do so.

Related