How to configure pytesseract to support text detection for non English language in windows 10?

Viewed 3816

I have tried pytesseract for English. It's working fine and generates expected result. But when it comes for other languages (eg: Arabic) other than english, it fails to do so and gives following error:

TesseractError: (1, 'Error opening data file C:\\Program Files (x86)\\Tesseract-OCR\\ara.traineddata 
Please make sure the TESSDATA_PREFIX environment variable is set to your "tessdata" directory. 
Failed loading language \'ara\' Tesseract couldn\'t load any languages! 
Could not initialize tesseract.')

Tried to get it (ara.traineddata) done from github, but can't get it done.

1 Answers

pytesseract is only wrapper on program tesseract (OCR developed by Google)

tesseract needs files with languages which you can find in its documentation: Data Files.

You can download ara.traineddata to some folder and run it with option --tessdata-dir some_folder and then it will use ara.traineddata from this folder.

If you save ara.traineddata in the same folder as you run code then you can use . (dot)

tesseract image.jpg stdout -l ara --tessdata-dir .

And the same you can do in pytesseract using config=

import pytesseract

text = pytesseract.image_to_string('image.jpg', lang='ara', config='--tessdata-dir .')

print(text)

Eventually you can use environment variable TESSDATA_PREFIX for this

import pytesseract
import os

os.environ['TESSDATA_PREFIX'] = '.'

text = pytesseract.image_to_string('text-ara.jpg', lang='ara')

print(text)

Later you can set TESSDATA_PREFIX directly in system or you may try to move ara.traineddata to folder with other files .traineddata. There should be somewhere eng.traineddata which you can try to find with programs/command like find


I tested it with this image which I found also in documentation: Command Line Usage

enter image description here


BTW: tesseract normally saves text in file but if you use stdout then it displays text in console.

Related