is text in image is bold or not?

Viewed 2027

I have been experimenting latley with Tesseract OCR. I am able to find characters in an image, but I have trouble finding only the bold characters in an image (know if a character in a document image is bold or not). I saw the function WordFontAttributes() mentioned in another question (Can I use OCR to detect font style (bold, italic)?) from the Tesseract API but I am not able to implement it in Python.

1 Answers

before it install tesseract 3.05 (4-th version doesn't support WordFontAttributes)

from tesserocr import PyTessBaseAPI, RIL, iterate_level


def get_words_info(image_path, tessdata_path):
    """
    get path to image and path to tessdata and return dict with info about each word
    """
    # api = PyTessBaseAPI(path=tessdata_path)
    with PyTessBaseAPI(path=tessdata_path) as api:
        api.SetImageFile(image_path)
        api.Recognize()
        iter = api.GetIterator()
        level = RIL.WORD

        result = []

        for r in iterate_level(iter, level):
            element = r.GetUTF8Text(level)
            word_attributes = r.WordFontAttributes()
            base_line = r.BoundingBox(level)

            if element:
                word_attributes['word'] = element
                word_attributes['position'] = base_line

            result.append(word_attributes)

        return result
Related