With regard to this question and this question, where I ask how to download thousands of PDF and processes them to extract their texts with OCR, I am hitting a brick wall again when it comes to enhancing the text outputs.
I am interested to extract texts of a bunch of PDF in order to search for surnames in the text (I do not need necessarily to be able to read the rest of the text). The PDF represent old newspaper articles, published between 1810 and 1832 and written in German Fraktur. This font seems to be particularly challenging for tesseract.
Q: How can I further improve the image quality for tesseract to - at least - have a change to find the surnames in the text? Which procedure would you suggest?
If we take this pdf as an example, I receive the following image when applying
convert -colorspace GRAY -resize 3000x -units PixelsPerInch example.pdf example-page.jpg
If I now use tesseract with
tesseract --tessdata-dir /usr/local/share/tessdata/ -l deu_frak example-page.jpg example-page.txt
it would perform terrible on that image with roughly 360 diacritics detected only. My text output is entirely scrambled.
When I use Fred's ImageMagick script textcleaner, applying either
textcleaner -g -e stretch -f 25 -o 10 -u -s 1 -T -p 10
or
textcleaner -g -e stretch -f 25 -o 20 -t 30 -u -s 1 -T -p 20
I get something like this
When I then run again tesseract with the above mentioned command, the resulting text is much better (around 700-800 diacritics detected) but still scrambled enough not to find most surnames of the text.
I know that the example page is a particular hard one, however, even pages, which are not inky prints and not skewed to begin with, yield mostly scrambled outputs and undecipherable surnames when processing them with tesseract and the above command.
For example this page
Q: How can I further improve the image quality for tesseract to - at least - have a change to find the surnames in the text? Which procedure would you suggest?
Edit: I do not know, whether training tesseract is needed or a good idea to deal with the given German Fraktur font, as GUI box editor seems to work reliably on MacOS, see for example, jTessBoxEditor, Qt-box-editor, or Tesseract-Box-Editor, nor did I understand how to train tesseract, see the tesseract training wiki here and another tutorial here.


