Tesseract in R does not recognize smaller fonts in the same document

Viewed 979

With beginner level expertise in R and asked to write codes to convert text from images into a txt file, I am struggling with the Tesseract and the Magick packages.

I will unfortunately not be able to upload the original document as it is confidential, but I have tried my best to replicate the same into a dummy image which is attached. The original is similarly structured as the attached example.

The document contains a line, which is in very small fonts. The code that I am running, reads most of the fonts correctly but does not read the fonts which are much smaller in size (around 6-6.5 font size or lesser in MS Word.)

This is a huge problem because the most vital piece of information, lies in those smaller fonts and not being able to read it, makes the whole exercise of conversion pretty much useless

I have followed 2 different versions of the code and both of them come with their on set of challenges:-

Version 1 -->

text5 <- ocr("D:/abc/dummy.PNG")
cat(text5)
write.table(text5, "D:/abc/Outputs/dummy.txt", sep="\t")

Problem with version 1 --> The output is generated in a few seconds, everything is just about perfect, but the text of input in smaller fonts is not at all acceptable.

Version 2 -->

test2 <- image_read("D:/abc/dummy.PNG") %>%
image_resize("3000") %>%
image_convert(colorspace = 'gray') %>%
image_trim() %>%
image_ocr()
cat(test2)
write.table(test2, "D:/abc/Outputs/dummy.txt", sep="\t")

Problem with Version 2 --> The output is slightly better, but still there is a lot of scope of improvement.

I tried multiple resources like source1, source2 and feel it has something to do with the low dpi of that particular line, but I am not sure how to go about it. I might be totally wrong so feel free to correct.

Optimistic to get some help from this forum![Sample image, original very similar to this but confidential, hence the sample]3

1 Answers

Does the document have the same format every time, or does it change?

If it were the same, you can just crop the place youre struggling with, then resize it little by little while using morphology operations, such as opening. What this does is, for every time you make the image larger, and unwanted white pixels appear inbetween your letters, it fills them again with the black pixels.

http://www.fmwconcepts.com/imagemagick/morphology/index.php

Edit: Added new comments.

Related