How to specify a specific English font for Tesseract v5 OCR?

Viewed 93

Problem

I am using Tesseract v5 and IronOCR to perform OCR on files written in the IBM Plex Mono font. It is having issues recognizing dotted zeroes because it is confusing them with 6, 8, and 9.

How can I train Tesseract OCR to recognize a dotted zero, or specify that the font is specifically IBM Plex Mono?

Here is a sample image. It is currently parsed as follows, with incorrectly parsed values highlighted in bold:

Input

Source Image

Output

X Y Z
0.3 0.0 0.0
1.8 0.0 0.0
3.8 0.3 06.06
1.1 1.2 0.0
06.9 0.8 0.0
3.0 3.1 06.0
1.7 0.6 0.0

Source Code

I am using IronOCR with Tesseract to parse. Here is my configuration for the parser:

Input.AddPdf("myfile.pdf");
Input.Deskew();  // fixes rotation and perspective
Input.DeNoise(); // fixes digital noise and poor scanning
Ocr.Configuration.BlackListCharacters = "X@©®¢*%,";
Ocr.Language = OcrLanguage.EnglishBest;
0 Answers
Related