How to trick camelot into reading from a byte-string rather than a real file

Viewed 204

I was wondering if it was possible to read a pdf into camelot not by giving it the path of a file, but rather a binary string containing the PDF data.

The reason I want to do this is that I have PDFs inside a zip-File, and rather than extracting the content into a temporary directory I would like to pass the byte data directly into camelot.

So far I have tried the following:

import ZipFile
import Path
from io import BytesIO
import camelot
 
zipFileName = Path("file.zip") # containing the PDF file
pdf = ZipFile(zipFileName).read("path_to_zip.zip")
# pdf now contains the content of the PDF and starts like this:
# b'%PDF-1.4\n%\x80\x81\x82\...
f = BytesIO(pdf)
tables = camelot(f)

This will lead to an error messages

File ~/opt/anaconda3/envs/brd-test/lib/python3.9/site-packages/camelot/handlers.py:41, in PDFHandler.__init__(self, filepath, pages, password)
     39     filepath = download_url(filepath)
     40 self.filepath = filepath
---> 41 if not filepath.lower().endswith(".pdf"):
     42     raise NotImplementedError("File format not supported")
     44 if password is None:

AttributeError: '_io.BytesIO' object has no attribute 'lower'

Obviously camelot really wants to operate on a file and checks if the extension matches ".pdf".

Any suggestion how to trick camelot into accepting the content rather than the file path?

1 Answers
Related