PDF parsing: using pdfminer and pandas

Viewed 1300

I am trying to parse a pdf file into csv format. In the pdf, there is a table without frame, so the method suggested here does not work. My idea is to use pdfminer to analyze the layout of the pdf, locate all textlines, and match the bbox location of each textlines to reconstruct the table.

So far I have successfully sorted the text lines into "left" and "right" column by comparing the x0 coordinates of each textline objects, and I am going to matching left and right lines based on their y0 coordinates. When I was trying to put the content of each lines into a pandas DataFrame, I got a TypeErrorL cannot concatenate a non-NDFrame object. Please help.

My code is as follow:

testfile = 'file location'
page_layouts = extract_layout_by_page(testfile)
l_lines = []
r_lines = []
for elem in page_layouts[0]:
    if isinstance(elem, pdfminer.layout.LTTextBoxHorizontal):
       for l in elem:
           (x0,y0,x1,y1) = l.bbox
           if x0 <= 65.35 and x0 >=65.33:
               l_lines.append(l)
           elif x0 <= 280.1 and x0 >= 279.9:
               r_lines.append(l)

csv = pd.DataFrame()
csv['l'] = 0
csv['r'] = 0

for i in r_lines:
    x = i.get_text().encode('ascii','ignore')
    csv['r'].append(x) 

Thank you in advance.

1 Answers

It's old but maybe it will help someone. I Guess that your mistake is adding text to a data frame. You should put this text in a series. You do something like this is the latest part:

l=[]
for i in r_lines:
    x = i.get_text().encode('ascii','ignore')
    l.append(x) 
df=pd.DataFrame()

#if you want to append vertically (adding a column):
df.append(l)
#if you want to append horizontally (adding a row):
s=pd.Series(l)
s.name ("Series") #you may custom format this to create a new name for each row added, with a counter for example. Giving a name is necessary when the series is appended to the df, else it raises a TypeError (unless you use ignore_index = True
df.append(s)
Related