Can I train tesseract using image/text pairs where the images and texts are just single words each? Most training examples I've seen on Github use lines, each line of text being an image with the correct text for that line. However I have a system which is already going to be producing word image/text pairs and I'd like to feed that back into training. Any reason why not? I know that there are page segmentation modes and that word segmentation and line segmentation are not the same thing. But I understand that psm only applies to inference and not training?
Update: I've posted this to the Tesseract github issues and the google group with no response there either. I'm not sure whether the question is badly formulated, or if it's just the case that noone knows the answer? I'm hoping that a bounty might encourage some input.