Is there a way to make pytesseract.image_to_pdf_or_hocr output both pdf and text data?
Currently I am doing like this:
pdf = pytesseract.image_to_pdf_or_hocr(fp.name, extension='pdf')
text = pytesseract.image_to_string(fp.name)
is there a way to do something like this so that tesseract runs only once? If no, what's a better way to do this?
pdf, text = pytesseract.image_to_pdf_or_hocr(fp.name, extension='pdf')