Optimising pdfminer

Viewed 2334

I am trying to use pdfminer.six in a production context to extract the text from a pdf. At the moment, for my benchmark 44 page document, it is taking approximately 18 seconds. I would like to reduce this as much as possible.

So far I have managed to reduce the time by 3 seconds, by turning caching = False. Does anyone have suggestions for how I can optimise this further? As far as I can tell using a module like multiprocessing to process the pages in parallel would not work because the underlying methods/functions are not abled to be pickled.

from pdfminer.pdfinterp import PDFResourceManager, PDFPageInterpreter
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
from pdfminer.pdfpage import PDFPage

path = "PATH/TO/MYPDF.pdf"
rsrcmgr = PDFResourceManager()
retstr = io.StringIO()
codec = 'utf-8'
laparams = LAParams()
device = TextConverter(rsrcmgr, retstr, codec=codec, laparams=laparams, showpageno= True)
fp = open(path, 'rb')
interpreter = PDFPageInterpreter(rsrcmgr, device)
password = ""
maxpages = None
caching = False
pagenos=set()

for page in PDFPage.get_pages(fp, pagenos, maxpages=maxpages, password=password,caching=caching, check_extractable=True):
    interpreter.process_page(page)

text = retstr.getvalue()
fp.close()
device.close()
retstr.close()
1 Answers

I am using pdfminer on python 3.8. I have an application that manipulates contents of a pdf document and though it is quite a chore to assemble words/tokens and determine where they occur in a tabular document, I had this all running fine in python 2.7, but moving to py3 and the latest version of pdfminer, it ran so slow that it was un acceptable. So after a lot of digging and profiling my code I found that because all of the print statements from the older version have been converted to log statements and the loggers created by pdfminer modules are all default to level.DEBUG, and since I have assigned handlers against the root logger to write log messages to a file the overall speed was highly impacted. Buy adding the following code after import of pdfminer modules and before instantiating any of the classes or calling them it now runs acceptably fast.

# set all pdfminer logging to WARN
pdflogs = [logging.getLogger(name) for name in logging.root.manager.loggerDict if name.startswith('pdfminer')]
for ll in pdflogs:
    ll.setLevel(logging.WARNING)
Related