I am doing OCR to read the 9-digit numbers. The 2 first images are the output of the segmentation model. Actually, they are grayscale. The 3rd is the input.
Currently, I use K-means with K=9 to group contours. It worked superbly on 1st case but did not in 2nd case. To deal with it, I set K=10. If the sum of distances decreases significantly, then remove either the first or the last box. Else, I keep K=9 and take all boxes. The result was better, but not really efficient in some cases (3rd). As you see in the 3rd case, the noise is far from the number and bigger than one digit.
How can I do cluster more efficiently? Thanks in advance!


