Extract table data from PDF that is formatted as picture

Viewed 862

I am trying to extract the data in the tables that start on p.52 of this document (a report from FAA).

The problem is that the tables are included as pictures. Any chance I can get some pointers on how to do that without doing it manually?

I have tried converting it to text using Adobe's OCR function, and I have also tried using the extract_tables function in R's tabulized package.

I could of course do this manually, but it would be good to know if there is a more efficient way of doing it.

1 Answers

It's possible, however its accuracy depends on the image. I always use grayscale images. Here an example of available tools. In your case, I'd suggest you take some screenshots of the tables and use the OCRFeeder to compare the results from GOCR and Tesseract.

sudo apt-get install gocr tesseract-ocr ocrfeeder

ocrfeeder -i image.jpg

After some manual checks, you can import this file in LibreOffice Calc, save it as 'csv', and import in R.

Related