How to extract text from table in an image and arrange them into spreadsheet OpenCV

Viewed 156

Edit added: Table format 1 , Table format 2 ,Table format 3 , stackoverflow only allow me to maximum 6 images. Actually there are 2 more but not so different just the table position are different. All the drawing are partially removed after processing, these are the original image with drawing and tables

Original Image and Image With Table only, the table only image is acquired after processing the Original Image, I would like to know how to use Pytesseract to extract text and arrange them properly into a spreadsheet , for now I am only able to extract them blindly using pytesseract.image_to_string and split them using split() and then save each splitted text into an array, which is not accurate at all because all data is arrange blindly.

#Get the string from the table only image with image_to_string()
data = pytesseract.image_to_string(table_only_image)
#Split text in data with split()
data_list = data.split()

#----------------------------Write Extracted Data Into Spreadsheet-------------------------

len_data = len(data_list)
cellcol = 1
workbook = Workbook()
sheet = workbook.active
sheet.title = "Tabulardata"

for cellrow in range(1,len_data):
    #Write data blindly without arrangement into the spreadsheet.
    sheet.cell(row=cellrow,column=cellcol).value = data_list[cellrow-1] 
0 Answers
Related