Training Tesseract: how to handle multiple whitespace characters in training images

Viewed 82

I am finetuning tesseract. I am using tesstrain. When creating the ground truth files, how should I handle multiple whitespace characters?

a.tiff enter image description here

There are two whitespaces between "(f)" and "any". I have no idea how many are between "30" and "(f)". Should I:

  1. Attempt to provide the correct number of whitespaces? (in a.gt.txt)
  2. Just use 1 whitespace between words always?
  3. It doesn't make a difference.

I've seen on other tangentially related questions that when doing the inference, there is an option to preserve interword space. Maybe that makes a difference:

0 Answers
Related