Image_to_string not reading text from tiff or tif files using pytesseract

Viewed 784

I am trying to read text from tif or tiff image files. These files have multiple pages.

When I print the array i only get true and then no text. However when i use .png files i am able to print the text.

Below is my code.

from PIL import Image, ImageSequence
import pytesseract
from pytesseract import image_to_string
import numpy as np
import cv2
test = Image.open(r'C:\Python\BG36820V1.tiff')
#test1 = Image.open(r'C:\Users\Documents\declaration.png')
testarray = np.array(test)
print(testarray)
print(pytesseract.image_to_string(Image.fromarray(testarray))

This is the out put for the test file:

[[ True  True  True ...  True  True  True]
 [ True  True  True ...  True  True  True]
 [ True  True  True ...  True  True  True]
 ...
 [ True  True  True ...  True  True  True]
 [ True  True  True ...  True  True  True]
 [ True  True  True ...  True  True  True]]

However this works fine with the test1.

[[[242 242 242 255]
  [242 242 242 255]
  [242 242 242 255]
  ...
  [242 242 242 255]
  [242 242 242 255]
  [242 242 242 255]]

 [[182 180 182 255]
  [182 180 182 255]
  [182 180 182 255]
  ...
  [182 180 182 255]
  [182 180 182 255]
  [182 180 182 255]]
g Request 4042337300021 submitted sucessfully

x
TYPE

i tried opencv to read tiff files i get format not supported.

How do i get to print the text from the tiff or tif files.

Any suggestions?

Regards, Ren.

1 Answers

Modified my entire code and converted the tiff files to jpeg file and it was able to read the text.

from PIL import Image, ImageSequence
import pytesseract
from pytesseract import image_to_string
import numpy as np
import cv2
import os
yourpath = r'C:\Python\'
for root, dirs, files in os.walk(yourpath, topdown=False):
    for name in files:
        print(os.path.join(root, name))
        if os.path.splitext(os.path.join(root, name))[1].lower() == ".tiff":
            if os.path.isfile(os.path.splitext(os.path.join(root, name))[0] + ".jpg"):
                print ("A jpeg file already exists for %s" % name)
            # If a jpeg is *NOT* present, create one from the tiff.
            else:
                outfile = os.path.splitext(os.path.join(root, name))[0] + ".jpg"
                try:
                    im = Image.open(os.path.join(root, name))
                    print ("Generating jpeg for %s" % name)
                    im.thumbnail(im.size)
                    im.save(outfile, "JPEG", quality=100)
                except Exception as e:
                    print (e)

test = Image.open(r'C:\Python\BG96254V1.jpeg')

testarray = np.array(test)
print(testarray)
print(pytesseract.image_to_string(Image.fromarray(testarray)))

This only reading the 1st page not the list of pages. Any suggestions how to make that change to read all the pages.

Thanks.

Related