Improve quality of image for tesseract OCR

Viewed 1389

With regard to this question and this question, where I ask how to download thousands of PDF and processes them to extract their texts with OCR, I am hitting a brick wall again when it comes to enhancing the text outputs.

I am interested to extract texts of a bunch of PDF in order to search for surnames in the text (I do not need necessarily to be able to read the rest of the text). The PDF represent old newspaper articles, published between 1810 and 1832 and written in German Fraktur. This font seems to be particularly challenging for tesseract.

Q: How can I further improve the image quality for tesseract to - at least - have a change to find the surnames in the text? Which procedure would you suggest?

If we take this pdf as an example, I receive the following image when applying

convert -colorspace GRAY -resize 3000x -units PixelsPerInch example.pdf example-page.jpg

enter image description here

If I now use tesseract with

tesseract --tessdata-dir /usr/local/share/tessdata/ -l deu_frak example-page.jpg example-page.txt

it would perform terrible on that image with roughly 360 diacritics detected only. My text output is entirely scrambled.

When I use Fred's ImageMagick script textcleaner, applying either

textcleaner -g -e stretch -f 25 -o 10 -u -s 1 -T -p 10

or

textcleaner -g -e stretch -f 25 -o 20 -t 30 -u -s 1 -T -p 20

I get something like this

enter image description here

When I then run again tesseract with the above mentioned command, the resulting text is much better (around 700-800 diacritics detected) but still scrambled enough not to find most surnames of the text.

I know that the example page is a particular hard one, however, even pages, which are not inky prints and not skewed to begin with, yield mostly scrambled outputs and undecipherable surnames when processing them with tesseract and the above command.

For example this page

enter image description here

Q: How can I further improve the image quality for tesseract to - at least - have a change to find the surnames in the text? Which procedure would you suggest?

Edit: I do not know, whether training tesseract is needed or a good idea to deal with the given German Fraktur font, as GUI box editor seems to work reliably on MacOS, see for example, jTessBoxEditor, Qt-box-editor, or Tesseract-Box-Editor, nor did I understand how to train tesseract, see the tesseract training wiki here and another tutorial here.

1 Answers
Related