how to increase resolution of text in scanned images in python?

Viewed 5116

I use tesseract-OCR to extract text from scanned images, For few images text is not properly recognized due to low resolution and output produced is some irrelevant characters.

Techniques applied:

  1. Increase the dpi to 300.

  2. Image pre- processing techniques in opencv.

  3. Upscaling of images using dnn_superres in opencv

  4. Noise removal techniques.

  5. Refereed git repos where super-resolution algorithm model is developed using Deep learning.

  6. Improve tesseract-ocr quality by training tessdata.

Reference Links:

  1. Improve OCR accuracy from scanned documents
  2. image processing to improve tesseract OCR accuracy

Sample Image:

enter image description here

Is there any simple way in python to improve the text without using any Deep learning model.

1 Answers

I am aware you would prefer to upscale these input images with using deep learning, but I would highly recommend experimenting with https://github.com/alexjc/neural-enhance, assuming you have the appropriate hardware to run the neural networks and deep learning.

The results for your OCR input images could be promising. The documentation for the code is quite substantial.

Hope this helps you!

Related