PYTHON - Extract specific table from word document of tables, and display in excel

Viewed 40

I'm trying to extract the 8th table from 500 word documents, and then put them into one page on an excel document, retaining the columns.

import pandas as pd
from docx.api import Document

path = 'C:\\Users\\'
worddocs_list = []
for filename in os.listdir(path):
    wordDoc = Document(path+"\\"+filename)
    worddocs_list.append(wordDoc)

for wordDoc in worddocs_list:
    table = document.tables[8]

    data = []

    for i, row in enumerate(table.rows):
        text = (cell.text for cell in row.cells)

        row_data = (text)
        data.append(row_data)
        print (data)

    df = pd.DataFrame(data)

print(df)

This gives an output of the pd.DataFrame, which is in the correct table format, however, it is only outputting the DataFrame from the first word document in the folder - how do I make it iterate through and pull all of the 8th tables?

The format it is giving is correct:

    1   2   3   4
0   5   6   7   8
1   9   10  11  12
2   13  14  15  16
3   17  18  19  20

etc.

How could I have it added to an excel document and iterate through the 500 other word documents to also add in the same format to excel?

All help is greatly appreciated, thank you.

0 Answers
Related