Extracting text from pdf containing multiple columns with pdfplumber (Python)

Viewed 1466

I am extracting texts from pdfs (with python) in order to analyze them so I am working a lot with scientific papers. I am using pdfplumber and it is working excellent the only problem is that such pdfs usually contain columns and I haven't found a way for my algorithm to recognize this.

My code is:

text = ""
with pdfplumber.open(r'example.pdf') as pdf:
    pages = pdf.pages
    for i, pg in enumerate(pages):
        text = text + " " + pages[i].extract_text(x_tolerance = 1)
text = text.replace('\n',' ')
text = text.replace('\r',' ')
text = text.replace('\no',' ')
text = text.replace('\nD',' ')
text = text.lower()
text = re.sub(r'[^a-zA-Z0-9\s]', ' ', text)

do you know a function that might help me? Thank you!

1 Answers

Assuming the columns are separated by " " then you can use the draw_rects() function to craft out the text from the columns.

The process requires you to detect the chars and spaces first using

im.reset().draw_rects(p0.chars)

And then using

text = p0.extract_text()

will give you the text format of the columns. This example illustrates how it can be done - Extract the column text from the PDF

Related