I'm training a model to read a specific type of captcha image. This is what the images look like:
They all have exactly 4 digits, which may or may not overlap slightly, and have random pastel colors. This url will generate a new random captcha on each refresh.
Before doing any OCR, I'm trying to clean up the image as best as I can. Here's my best solution so far:
import cv2
import numpy as np
from matplotlib import pyplot as plt
def blob_filter(input, min_size):
nb_components, output, stats, centroids = cv.connectedComponentsWithStats(input, connectivity = 8)
sizes = stats[1:, -1];
nb_components = nb_components - 1
filtered = np.zeros((output.shape))
for i in range(0, nb_components):
if sizes[i] >= min_size:
filtered[output == i + 1] = 255
return filtered
def read_and_preprocess(image_path):
original = cv2.imread(image_path)
hsv = cv2.cvtColor(original, cv2.COLOR_BGR2HSV)
_, _, v = cv2.split(hsv)
blurred = cv2.blur(v, (2,2))
_, thresh = cv2.threshold(blurred, 180, 255, cv2.THRESH_BINARY)
filtered = blob_filter(thresh, 80)
return np.asarray(filtered)
im_path = "yl5ux.jpg"
example = read_and_preprocess(im_path)
plt.imshow(example, cmap='gray')
I noticed that the digits are most visible in the V channel in HSV space, so I basically just blur and threshold this channel, then remove most of the blob artefacts that are left. This produces something like this:
This result is generally satisfactory but the edges are still very noisy and the thresholds some times cut off parts of some digits.
I'm wondering if anybody can suggest a different approach or additional steps to improve the legibility of these digits.
I've tried some morphological operations, low-pass filtering in frequency domain and a low-variance filter, but none of them produced better or robust results for all images.
This link provides 2000 sample captchas with labels for download, for anybody interested in experimenting.

