Pytesseract does not recognise certain characters even after image pre-proccessing and training for new font

Viewed 38

I tried getting the text from the image below using pytesseract. The image

If i use the default traineddata (eng.traineddata) i get the following result: "AA2 + b2 = cA2. Who made this formula for a triangle?"

After training tesseract for a new font (Droid Serif Bold) and using the result traineddata i get this result: "AA2 + bA2 = cA2. Who made this formula for a triangle?"

What else can i do to get close to 99% accuracy? I think 99% is a fair expectation since the image is very clear, little to no image noise, high contrast between font color and backround color.

0 Answers
Related