Reading Tables from PDFs in S3 bucket using Camelot or Tabula packages: s3 URL

Viewed 100

Can Python packages that pull tables from PDFs, such as Tabula and Camelot, read in the PDF from an S3 bucket - like with Pandas. For example, I can read a CSV file from the S3 bucket like this:

df = pd.read_pdf("s3://us-east-1-name/Test/Testfile.csv")

I want to be able to do the same thing using Tabula or Camelot:

dfs = tabula.read_pdf("s3://us-east-1-name/Test/Testfile.pdf", pages='all')

tables = camelot.read_pdf("s3://us-east-1-name/Test/Testfile.pdf")

I get an "HTTP Error 403: Forbidden" or "[Errno 2] No such file or directory." But there is no issue with the S3 locations. Does anyone know how I can pass an S3 URL/API with Tabula or Camelot.

0 Answers
Related